Visual Search OpenVINO Models
This repository contains the four OpenVINO FP16 model packages used by the Visual Search Windows desktop application. They are runtime artifacts, not newly trained models. Each package preserves its upstream attribution and includes the metadata needed by the application.
Included models
| Directory | Purpose | Upstream model | License |
|---|---|---|---|
siglip2-so400m-patch16-naflex-openvino-fp16 |
Multilingual global and region image/text embeddings | google/siglip2-so400m-patch16-naflex |
Apache-2.0 |
pp-ocrv6-medium-openvino-fp16 |
Text detection and multilingual recognition | PaddlePaddle/PP-OCRv6_medium_det_onnx and PaddlePaddle/PP-OCRv6_medium_rec_onnx |
Apache-2.0 |
yoloe-26s-pf-openvino-fp16 |
Prompt-free object evidence, small variant | Ultralytics YOLOE-26S-PF | AGPL-3.0 unless covered by an Ultralytics Enterprise license |
yoloe-26m-pf-openvino-fp16 |
Prompt-free object evidence, medium variant | Ultralytics YOLOE-26M-PF | AGPL-3.0 unless covered by an Ultralytics Enterprise license |
Because the bundle contains artifacts under different licenses, the repository metadata uses license: other. Redistribution and use remain subject to each upstream model's license and terms.
Runtime use
Visual Search downloads files from an immutable repository revision. Before installation, every file is checked against the expected byte size and SHA-256 digest in the application's models.json manifest. Downloads support resume and are moved into place only after validation succeeds.
- SigLIP2 uses a fixed 256-patch transformer input and explicit position embeddings for stable OpenVINO execution while preserving NaFlex aspect-ratio behavior.
- PP-OCRv6 includes separate detection and recognition OpenVINO IR files plus preprocessing and character-dictionary configuration.
- YOLOE uses a fixed prompt-free vocabulary. Its detections contribute soft ranking evidence and are not used as a hard search filter.
Qwen3.5 INT8 and INT4 are intentionally not duplicated here. Visual Search downloads those optional VLM packages from their official OpenVINO Hugging Face repositories.
Repository layout
siglip2-so400m-patch16-naflex-openvino-fp16/
pp-ocrv6-medium-openvino-fp16/
yoloe-26s-pf-openvino-fp16/
yoloe-26m-pf-openvino-fp16/
Package-level README.md, configuration, validation metadata, and SHA-256 manifests are included where applicable. The models are designed for the Visual Search runtime and are not presented as a standalone Python API.
Limitations
- Accuracy and latency depend on the Intel GPU, driver, image resolution, language, and search configuration.
- YOLOE uses a fixed vocabulary and may not cover every object category.
- OCR accuracy varies with text size, rotation, blur, and image quality.
- Review the upstream model cards for training data, intended uses, and additional limitations.
Attribution
The repository contains format conversions of upstream model weights. Refer to the upstream model cards and package-level metadata for source revisions and attribution.