Granite Embedding 97M Multilingual R2 โ€” Core ML (int8)

A Core ML conversion of ibm-granite/granite-embedding-97m-multilingual-r2 at revision 835ad14087e140460703cf0fae09f97d469d65c2, for on-device memory recall in the Character Agent iOS app. Not affiliated with or endorsed by IBM.

Changes from the original

  • Converted from PyTorch to an iOS 18 Core ML package (GraniteEmbedding97mMultilingualR2.mlpackage).
  • fp16 compute, with finite attention masks and fp32 rotary tables.
  • Enumerated input lengths: 16, 32, 64, 128, 256, 512 tokens.
  • Weights quantized to int8 (98 MB).

Inputs and output

  • input_ids and attention_mask (Int32, one of the enumerated lengths).
  • embedding: the sentence embedding (384 dimensions), L2-normalized.
  • Tokenize with the original tokenizer.json from the base model at the revision above (byte-level BPE).

Parity with the original

Cosine similarity to the PyTorch model on 2,374 texts in 13 languages: minimum 0.99924, mean 0.99979 (Neural Engine and CPU).

Files and SHA-256

File Bytes SHA-256
GraniteEmbedding97mMultilingualR2.mlpackage/Manifest.json 617 4b1d83a52239e83bed576458e8d3e1ba4aa303385c619e825ebd62f1ff065032
GraniteEmbedding97mMultilingualR2.mlpackage/Data/com.apple.CoreML/model.mlmodel 172,838 c0553e1cd2455ac6c5196d8658768dcc9e77ced18cf5c847e002af6951b125a7
GraniteEmbedding97mMultilingualR2.mlpackage/Data/com.apple.CoreML/weights/weight.bin 98,079,232 a9ff4dce357172570f18de7062b2a06f88d6b9b439f38d99ae110af36a810b1c

License

Apache License 2.0, the same as the base model (see LICENSE). Base model ยฉ IBM.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for characteragent/granite-embedding-97m-multilingual-r2-coreml

Quantized
(18)
this model