Florence-2-base-ft vision encoder — Core ML (fp16)
The vision encoder (DaViT) of Florence-2-base-ft, converted to Core ML so the image side of captioning runs on Apple GPUs. Built for Retriever, a local semantic search over personal video archives on macOS.
What was modified: format only. The ONNX export from
onnx-community/Florence-2-base-ft (vision_encoder.onnx, fp32) was
simplified with fixed input shape (onnx-simplifier), loaded with
onnx2torch, wrapped with ImageNet normalization and converted with
coremltools 9 to an ML Program with float16 weights.
Input and output
Input image |
768×768 RGB image, values 0–255 (the model scales to [0, 1] and applies mean [0.485, 0.456, 0.406], std [0.229, 0.224, 0.225] itself) |
Output embedding |
image_features, shape (1, 577, 768), float16; feed it to the text encoder exactly like the ONNX output |
Fidelity against the ONNX original
40 real frames, per-token cosine against the ONNX features and the generated caption compared verbatim:
| Compute unit | ms / image | Min cosine | Identical captions |
|---|---|---|---|
| GPU | 128 | 0.9933 | 38 / 40 |
| Neural Engine | 1189 (compilation fails, falls back) | 0.9660 | 35 / 40 |
| CPU (Core ML) | 188 | 0.2808 | 8 / 40 |
ONNX Runtime on CPU: 744 ms per image. Run it on the GPU
(MLComputeUnits.cpuAndGPU); the Neural Engine cannot compile the full
graph and the Core ML CPU path is numerically off.
License
Florence-2 is released by Microsoft under the MIT license; this conversion keeps it.
- Downloads last month
- 7
Model tree for antonlnz/florence2-base-vision-coreml
Base model
microsoft/Florence-2-base-ft