IP-Adapter Plus (SD 1.5) for phones β ONNX, CPU
The CPU half of IP-Adapter for Nightmare Mobile's SD 1.5 Swap v2 models (checkpoints converted with npuforge as SD1.5 Swap). A Swap v2 UNet runs on the Snapdragon NPU and takes each cross-attention layer's image-prompt K and V as inputs; these files turn a reference picture into those K/V on the phone's CPU, once per picture:
picture β clip_vit_h (penultimate hidden state) β ip_<adapter>_head β ipk_0..15, ipv_0..15
| file | what | from |
|---|---|---|
clip_vit_h_w16qdq.onnx |
CLIP ViT-H/14 image encoder, truncated to 31 of 32 layers (the penultimate hidden state IP-Adapter Plus reads); int16 weights per output channel behind DequantizeLinear (opset 21), fp32 compute |
h94/IP-Adapter models/image_encoder (OpenCLIP ViT-H-14 laion2B, MIT) |
ip_plus_head_w8qdq.onnx |
IP-Adapter Plus: the Resampler (16 tokens) and each layer's to_k_ip / to_v_ip; int8 weights per output channel, fp32 compute |
h94/IP-Adapter models/ip-adapter-plus_sd15.safetensors |
ip_face_head_w8qdq.onnx |
IP-Adapter Plus Face, the same shape (int8 weights per output channel, fp32 compute) | h94/IP-Adapter models/ip-adapter-plus-face_sd15.safetensors |
IO: encoder pixels [1,3,224,224] (CLIP-normalised, centre-cropped) β hidden [1,257,1280];
head hidden β ipk_i [1,inner,16] and ipv_i [1,16,inner] for i = 0..15 in the Swap v2 UNet's
input order (down_blocks, up_blocks, mid_block). V comes out at scale 1: the app multiplies it by
the IP-Adapter strength. The unconditional side is the same head on an all-zeros pixel tensor.
Accuracy β worst cosine of any output K/V tensor vs the fp32 path, over 12 reference pictures
and the zeros image (check.json):
| encoder weights | Plus | Face | file |
|---|---|---|---|
| int16 per channel (shipped) | 0.99999995 | 0.99999991 | 1.17 GB |
| int8 per 64-block | 0.9925 | 0.9999 | 0.63 GB |
| int8 per channel | 0.991 | 0.9987 | 0.59 GB |
| int8 dynamic (activations too) | 0.838 | 0.976 | 0.59 GB |
ViT-H's activation outliers break int8 for Plus, whose Resampler amplifies small errors. The heads
are int8 (worst 0.99964 Plus, 0.99991 Face). Peak RSS running the int16 encoder: ~1.4 GB (fp16
weights behind Cast were as exact but ONNX Runtime expanded every weight up front: 3.3 GB).
Licence
Apache-2.0, as the IP-Adapter weights (Tencent AI Lab, h94/IP-Adapter); the image encoder is
OpenCLIP ViT-H-14 (laion2B), MIT. Converted to ONNX and quantized; no retraining.
SDXL heads
For npuforge's SDXL Swap template (70 cross-attention layers, ip_targets.json). Same shared encoder
(clip_vit_h_w16qdq.onnx, ViT-H/14 penultimate hidden state); one head per adapter: Resampler β 16 tokens β
each layer's to_k_ip / to_v_ip, outputs ipk_0..69 [1, inner, 16] and ipv_0..69 [1, 16, inner] at scale 1,
in the template's order. int8 weight-only (DequantizeLinear per output channel), fp32 compute.
| file | adapter | bytes | SHA-256 | worst K/V cosine vs fp32 (12 pictures + zeros image, through the int16 encoder) |
|---|---|---|---|---|
ip_plus_sdxl_head_w8qdq.onnx |
h94 ip-adapter-plus_sdxl_vit-h |
425119893 | f2aaf1150f81567aa925f9768781592ec4d2983a7b65dd511980e9fc69df02eb |
0.99978 |
ip_face_sdxl_head_w8qdq.onnx |
h94 ip-adapter-plus-face_sdxl_vit-h |
425119893 | 9cd9c207e47fe2eb024a8c7c5c1e10979a513a31f1697ec719bf4bc6ba3010bb |
0.99995 |
Source weights: h94/IP-Adapter (Apache-2.0).