Jina Code Embeddings 0.5B: Core ML W8A16

Compiled Core ML artifacts for code and natural-language embeddings in an 896-dimensional space, for code search, code-to-code retrieval, code-to-comment matching and completion lookup.

Requirements

  • Apple Silicon, macOS 15 or iOS 18 or later (the declared deployment target).
  • The .mlmodelc artifacts are precompiled, so they load without an on-device compile step.
  • A host implementation of the tokenisation, function selection and prompting described by manifest.json and the model's metadata.json.

Download

hf download 1of2/jina-code-embeddings-0.5b-coreml-w8 --local-dir ./model

Keep the directory structure intact; manifest.json is at the root. Pin a Hub revision in production.

Runtime contract

  • Precision: W8A16 (int8-compressed weights, float16 activations).
  • Output: 896-dimensional embeddings with last-token pooling; L2-normalise before use.
  • Matryoshka dimensions: 64, 128, 256, 512 and 896. Truncate, then normalise again.
  • Model: text_multifunc.mlmodelc, functions bucket_<n> for sequence lengths 32, 64, 128, 256, 512, 1,024 and 2,048, plus batch-4 functions bucket_<n>_b4 at 32, 64 and 128.
  • Inputs: input_ids, position_ids and selector (a one-hot over the last real token). The tower is causal and takes no attention mask. Pad with <|endoftext|> (id 151643) to the smallest bucket that fits.

Task prompts

Prefix queries and documents with the pair for the task, as recorded in manifest.json under taskPrompts:

Task Use
nl2code (default) natural-language query to code
code2code code to equivalent code
code2nl code to comment or docstring
code2completion start of a snippet to its completion
qa technical question to answer

Load a function

import CoreML

let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine
config.functionName = "bucket_128"
let model = try MLModel(
    contentsOf: URL(fileURLWithPath: "./model/text_multifunc.mlmodelc"),
    configuration: config
)
print(model.modelDescription)

A compute-unit preference is not proof of exclusive Neural Engine execution; placement depends on the function, hardware, OS and runtime.

Compatibility

The manifest's spaceID is jinaai/jina-code-embeddings-0.5b:896:w8a16. Re-embed existing indexes when changing space IDs, precision variants, prompts or dimensions. Do not mix W8A16 and W16A16 vectors in one index without evaluating it.

Validation

Checked against fp32 references from the source model with cosine-similarity gates.

License and attribution

Derived from Jina AI's jina-code-embeddings-0.5b at revision 4db235132dafbe56a8b9c5f59b59795ecf58a4a7. The weights are distributed under CC BY-NC 4.0: attribution is required and use is non-commercial. Commercial use needs a licence from Jina AI. This derivative converts the model to Core ML W8A16. It is not an official Jina AI release.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 1of2/jina-code-embeddings-0.5b-coreml-w8

Finetuned
(4)
this model