Comic region CLIP for Core ML
This is an on-device, fixed-label classifier for ColorComic's semantic coloring hints. It uses OpenAI CLIP ViT-B/32 vision weights and twelve ColorComic text prompts. It is not trained or evaluated on manga labels. Its label scores are zero-shot similarities; the synthetic numerical checks below establish conversion fidelity, not manga accuracy.
Interface
RegionClassifier.mlpackage has input image Float32 [1,3,224,224] and output probabilities Float32 [1,12]. Input is RGB, resized with CLIP's short edge 224 and center crop 224, scaled to [0,1], then normalized by mean [0.48145466, 0.4578275, 0.40821073] and standard deviation [0.26862954, 0.26130258, 0.27577711]. The caller supplies this normalized tensor; the package does no image preprocessing. Output labels follow labels.json, in the original ColorComic order.
The conversion encodes the same text prompts with the checkpoint's text tower, L2-normalizes each text prototype, L2-normalizes the projected vision embedding, multiplies dot products by 100, and applies softmax. Only the vision tower and fixed prototypes are stored in the package. Weights use Core ML FP16 conversion; input and output tensors remain Float32. The deployment target is iOS 16.
Reproduce
Source: openai/clip-vit-base-patch32 at commit 3d74acf9a28c67741b2f4f2ea7635f0aaf6f0268. Source pytorch_model.bin: 605,247,071 bytes, SHA-256 a63082132ba4f97a80bea76823f544493bffa8082296d62d71581a4feff1576f.
On macOS with Python 3.11 and the pinned versions in requirements.txt:
hf download openai/clip-vit-base-patch32 pytorch_model.bin config.json vocab.json merges.txt tokenizer_config.json special_tokens_map.json preprocessor_config.json --revision 3d74acf9a28c67741b2f4f2ea7635f0aaf6f0268 --local-dir .source
python convert.py
python validate.py
python manifest.py
The converter depends on the Hugging Face Transformers implementation of CLIP. convert.py contains the complete fixed prompt list and the classifier operation. validation.json records three deterministic synthetic cases against PyTorch on both CPU only and CPU/GPU Core ML execution. All six kept the same top label; maximum absolute probability difference was 0.01620 on CPU and 0.00164 on CPU/GPU. manifest.json records each package file's size and SHA-256. This repository contains no source PyTorch weights, only the derived Core ML package and twelve text prototypes.
Attribution and license
Conversion scripts and metadata: MIT, Copyright (c) 2026 VioletXF (LICENSE). CLIP model and weights: MIT, Copyright (c) 2021 OpenAI (LICENSE-OpenAI-CLIP). The twelve prompts and ordering are from ColorComic commit 8521e525963893b69649c4cf2a747095928916ed, MIT, Copyright (c) 2026 Vikas Tiwari (LICENSE-ColorComic).
- Downloads last month
- 9