IP-Adapter Plus (SD 1.5) for phones β€” ONNX, CPU

The CPU half of IP-Adapter for Nightmare Mobile's SD 1.5 Swap v2 models (checkpoints converted with npuforge as SD1.5 Swap). A Swap v2 UNet runs on the Snapdragon NPU and takes each cross-attention layer's image-prompt K and V as inputs; these files turn a reference picture into those K/V on the phone's CPU, once per picture:

picture β†’ clip_vit_h (penultimate hidden state) β†’ ip_<adapter>_head β†’ ipk_0..15, ipv_0..15
file what from
clip_vit_h_w16qdq.onnx CLIP ViT-H/14 image encoder, truncated to 31 of 32 layers (the penultimate hidden state IP-Adapter Plus reads); int16 weights per output channel behind DequantizeLinear (opset 21), fp32 compute h94/IP-Adapter models/image_encoder (OpenCLIP ViT-H-14 laion2B, MIT)
ip_plus_head_w8qdq.onnx IP-Adapter Plus: the Resampler (16 tokens) and each layer's to_k_ip / to_v_ip; int8 weights per output channel, fp32 compute h94/IP-Adapter models/ip-adapter-plus_sd15.safetensors
ip_face_head_w8qdq.onnx IP-Adapter Plus Face, the same shape (int8 weights per output channel, fp32 compute) h94/IP-Adapter models/ip-adapter-plus-face_sd15.safetensors

IO: encoder pixels [1,3,224,224] (CLIP-normalised, centre-cropped) β†’ hidden [1,257,1280]; head hidden β†’ ipk_i [1,inner,16] and ipv_i [1,16,inner] for i = 0..15 in the Swap v2 UNet's input order (down_blocks, up_blocks, mid_block). V comes out at scale 1: the app multiplies it by the IP-Adapter strength. The unconditional side is the same head on an all-zeros pixel tensor.

Accuracy β€” worst cosine of any output K/V tensor vs the fp32 path, over 12 reference pictures and the zeros image (check.json):

encoder weights Plus Face file
int16 per channel (shipped) 0.99999995 0.99999991 1.17 GB
int8 per 64-block 0.9925 0.9999 0.63 GB
int8 per channel 0.991 0.9987 0.59 GB
int8 dynamic (activations too) 0.838 0.976 0.59 GB

ViT-H's activation outliers break int8 for Plus, whose Resampler amplifies small errors. The heads are int8 (worst 0.99964 Plus, 0.99991 Face). Peak RSS running the int16 encoder: ~1.4 GB (fp16 weights behind Cast were as exact but ONNX Runtime expanded every weight up front: 3.3 GB).

Licence

Apache-2.0, as the IP-Adapter weights (Tencent AI Lab, h94/IP-Adapter); the image encoder is OpenCLIP ViT-H-14 (laion2B), MIT. Converted to ONNX and quantized; no retraining.

SDXL heads

For npuforge's SDXL Swap template (70 cross-attention layers, ip_targets.json). Same shared encoder (clip_vit_h_w16qdq.onnx, ViT-H/14 penultimate hidden state); one head per adapter: Resampler β†’ 16 tokens β†’ each layer's to_k_ip / to_v_ip, outputs ipk_0..69 [1, inner, 16] and ipv_0..69 [1, 16, inner] at scale 1, in the template's order. int8 weight-only (DequantizeLinear per output channel), fp32 compute.

file adapter bytes SHA-256 worst K/V cosine vs fp32 (12 pictures + zeros image, through the int16 encoder)
ip_plus_sdxl_head_w8qdq.onnx h94 ip-adapter-plus_sdxl_vit-h 425119893 f2aaf1150f81567aa925f9768781592ec4d2983a7b65dd511980e9fc69df02eb 0.99978
ip_face_sdxl_head_w8qdq.onnx h94 ip-adapter-plus-face_sdxl_vit-h 425119893 9cd9c207e47fe2eb024a8c7c5c1e10979a513a31f1697ec719bf4bc6ba3010bb 0.99995

Source weights: h94/IP-Adapter (Apache-2.0).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support