CLIP_ViT-B32

CLIP ViT-B/32 embeds images and text into the same representation space, allowing semantic similarity, text-to-image retrieval, image-to-text retrieval and cross-modal exploration.

This ONNX package was produced by EIDORA from the upstream model listed below. EIDORA is not the original model creator.

Recommended Use

  • Text-to-image and image-to-text retrieval.
  • Semantic exploration of image collections using natural-language queries.
  • Comparing text and images in the same embedding space.
  • Cross-modal clustering and similarity experiments.

Limitations

  • Treating similarity scores as evidence of identity, authorship, intent or sensitive attributes.
  • Collections where language, geographic or cultural bias has not been evaluated.
  • Applications requiring domain-specific semantic accuracy without testing on the target collection.

CLIP was trained on a large internet-derived image-text corpus and can reflect biases and gaps in that data. Its upstream model card recommends task-specific testing before deployment. Cultural and historical collections can differ substantially from its training distribution, so similarity results should be evaluated in the context of the target collection.

Compute Tier

medium — Provisional until EIDORA CPU benchmarks are recorded.

Input

Images

  • Modality: image
  • Color space: RGB
  • Tensor layout: NCHW
  • Model input: 224 x 224

Text

  • Modality: text
  • Maximum tokens: 77
  • Token IDs type: int64
  • Attention-mask type: int64

Output

  • image output: embedding with shape [batch, 512]
  • text output: embedding with shape [batch, 512]
  • Type: float32
  • Normalized: true
  • Similarity: cosine

Preprocessing

Image

  • Resize mode: resize_shorter_edge_then_center_crop
  • Resize size: 224
  • Interpolation: bicubic
  • Rescale: 1/255
  • Mean: [0.48145466, 0.4578275, 0.40821073]
  • Std: [0.26862954, 0.26130258, 0.27577711]
  • Mean/std normalization inside ONNX: true

Text

  • Tokenizer: CLIPTokenizer
  • Source: openai/clip-vit-base-patch32
  • Maximum length: 77
  • Padding: max_length
  • Truncation: true
  • Tokenization inside ONNX: false

Model Source and Architecture

  • Exact checkpoint: openai/clip-vit-base-patch32
  • Pinned revision: b97b0100e55e367c057773c2a614676470b0d575
  • Checkpoint source: https://huggingface.co/openai/clip-vit-base-patch32
  • Original model: CLIP ViT-B/32
  • Original implementation: https://github.com/openai/CLIP
  • Type: dual_encoder
  • Image encoder:
    • Architecture: Vision Transformer
    • Variant: ViT-B/32
    • Image size: 224
    • Patch size: 32
    • Hidden size: 768
  • Text encoder:
    • Architecture: Transformer
    • Hidden size: 512
    • Max positions: 77
  • Projection dimension: 512

Training Data and Checkpoint Provenance

Approximately 400 million image-text pairs collected from the internet. The exact training dataset was not released.

Upstream license note: The OpenAI CLIP source repository is MIT licensed. The repository does not provide a separate explicit license statement for the pretrained weights, so redistribution terms for the weights should be reviewed before public release.

Research Context

CLIP was introduced as a contrastive image-text representation model trained to associate images with natural-language descriptions. The original work evaluates zero-shot transfer across a broad set of computer-vision tasks. The upstream model card describes CLIP primarily as a research model and recommends evaluating its behavior and limitations for the specific collection and task in which it is used.

For EIDORA, the shared image-text embedding space makes this checkpoint useful for exploratory semantic search, cross-modal retrieval and comparison of visual material with natural-language queries. Results should be interpreted as representation similarity rather than evidence of identity, authorship, intent or other sensitive attributes.

Attribution and Licensing

Based on the OpenAI CLIP ViT-B/32 checkpoint. Review the upstream checkpoint and pretrained-weight redistribution terms before public release.

Citation

@InProceedings{pmlr-v139-radford21a,
  title = {Learning Transferable Visual Models From Natural Language Supervision},
  author = {Radford, Alec and Kim, Jong Wook and Hallacy, Chris and Ramesh, Aditya and Goh, Gabriel and Agarwal, Sandhini and Sastry, Girish and Askell, Amanda and Mishkin, Pamela and Clark, Jack and Krueger, Gretchen and Sutskever, Ilya},
  booktitle = {Proceedings of the 38th International Conference on Machine Learning},
  pages = {8748--8763},
  year = {2021},
  editor = {Meila, Marina and Zhang, Tong},
  volume = {139},
  series = {Proceedings of Machine Learning Research},
  month = {18--24 Jul},
  publisher = {PMLR},
  url = {https://proceedings.mlr.press/v139/radford21a.html}
}

Package Information

  • Package version: 0.1.0
  • ONNX opset: 17
  • Exporter: eidora-onnx 1.0.0
Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EIDORA/CLIP_ViT-B32

Quantized
(12)
this model

Collection including EIDORA/CLIP_ViT-B32