On-device search models for Apple silicon

Ready-to-load conversions of text, image and speech models for macOS on Apple silicon, gathered in one repo so they can be fetched with plain curl, with no Python, Xcode or conversion step on the Mac that uses them.

Nothing here is fine-tuned: every model is a format conversion or a copy of the upstream weights, listed below with its source and licence.

Path What Source Licence
qwen3-embedding-0.6b-coreai.aimodel, qwen3-embedding-0.6b-coreai-tokenizer.bin Text embedding, exported to Core AI (fp16, input 2–128 tokens, unpadded, output [1, 1024]), and a tokenizer table Qwen/Qwen3-Embedding-0.6B @ 97b0c614 Apache 2.0
siglip2-base-patch16-224-{image,text}.mlpackage, siglip2-base-patch16-224-tokenizer.bin Image–text embedding: the image and text towers as Core ML, and a text tokenizer table google/siglip2-base-patch16-224 @ 75de2d55 Apache 2.0
bge-reranker-v2-m3-mlx/ Cross-encoder reranker, MLX flaglow/BAAI-bge-reranker-v2-m3-mlx-fp16 @ a8ecc6ca, from BAAI/bge-reranker-v2-m3 Apache 2.0
whisper-large-v3-turbo-coreml/ Speech to text, Core ML compiled for the Neural Engine (632 MB palettized) argmaxinc/whisperkit-coreml @ 0f63a780, openai_whisper-large-v3-v20240930_turbo_632MB; tokenizer from openai/whisper-large-v3 @ 06f233fe MIT

The .bin tokenizer tables are a compact custom format derived from the upstream tokenizer.json; use the upstream tokenizer with any other runtime.

The Core AI model needs macOS 27.

Downloads last month
20
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support