Visual Search OpenVINO Models

This repository contains the four OpenVINO FP16 model packages used by the Visual Search Windows desktop application. They are runtime artifacts, not newly trained models. Each package preserves its upstream attribution and includes the metadata needed by the application.

Included models

Directory Purpose Upstream model License
siglip2-so400m-patch16-naflex-openvino-fp16 Multilingual global and region image/text embeddings google/siglip2-so400m-patch16-naflex Apache-2.0
pp-ocrv6-medium-openvino-fp16 Text detection and multilingual recognition PaddlePaddle/PP-OCRv6_medium_det_onnx and PaddlePaddle/PP-OCRv6_medium_rec_onnx Apache-2.0
yoloe-26s-pf-openvino-fp16 Prompt-free object evidence, small variant Ultralytics YOLOE-26S-PF AGPL-3.0 unless covered by an Ultralytics Enterprise license
yoloe-26m-pf-openvino-fp16 Prompt-free object evidence, medium variant Ultralytics YOLOE-26M-PF AGPL-3.0 unless covered by an Ultralytics Enterprise license

Because the bundle contains artifacts under different licenses, the repository metadata uses license: other. Redistribution and use remain subject to each upstream model's license and terms.

Runtime use

Visual Search downloads files from an immutable repository revision. Before installation, every file is checked against the expected byte size and SHA-256 digest in the application's models.json manifest. Downloads support resume and are moved into place only after validation succeeds.

  • SigLIP2 uses a fixed 256-patch transformer input and explicit position embeddings for stable OpenVINO execution while preserving NaFlex aspect-ratio behavior.
  • PP-OCRv6 includes separate detection and recognition OpenVINO IR files plus preprocessing and character-dictionary configuration.
  • YOLOE uses a fixed prompt-free vocabulary. Its detections contribute soft ranking evidence and are not used as a hard search filter.

Qwen3.5 INT8 and INT4 are intentionally not duplicated here. Visual Search downloads those optional VLM packages from their official OpenVINO Hugging Face repositories.

Repository layout

siglip2-so400m-patch16-naflex-openvino-fp16/
pp-ocrv6-medium-openvino-fp16/
yoloe-26s-pf-openvino-fp16/
yoloe-26m-pf-openvino-fp16/

Package-level README.md, configuration, validation metadata, and SHA-256 manifests are included where applicable. The models are designed for the Visual Search runtime and are not presented as a standalone Python API.

Limitations

  • Accuracy and latency depend on the Intel GPU, driver, image resolution, language, and search configuration.
  • YOLOE uses a fixed vocabulary and may not cover every object category.
  • OCR accuracy varies with text size, rotation, blur, and image quality.
  • Review the upstream model cards for training data, intended uses, and additional limitations.

Attribution

The repository contains format conversions of upstream model weights. Refer to the upstream model cards and package-level metadata for source revisions and attribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support