LFM2.5-1.2B-Instruct for ExecuTorch

ExecuTorch exports of LiquidAI/LFM2.5-1.2B-Instruct (revision 0f604ada3f76) for on-device inference with the openweights Android app or any ExecuTorch 1.4.0 runtime.

Files

Windows in this repo: XNNPACK (CPU) at 2k to 32k; MediaTek NeuroPilot MT6991 (Dimensity 9400) at 512 to 8k. The window is fixed inside the file: the runtime allocates the whole KV cache at load, so pick the largest window the device can hold (fits_phone_budget in each folder's config.json is the estimate against a 5 GB budget).

Backend Target File Window Size Smoke test
XNNPACK (CPU) any arm64 xnnpack/LFM2.5-1.2B-Instruct-8da4w-gptq-2k.pte 2,048 tokens 0.80 GB passed ("Paris") with a completion prompt (the chat template did not render)
XNNPACK (CPU) any arm64 xnnpack/LFM2.5-1.2B-Instruct-8da4w-gptq-4k.pte 4,096 tokens 0.80 GB passed ("Paris") with a completion prompt (the chat template did not render)
XNNPACK (CPU) any arm64 xnnpack/LFM2.5-1.2B-Instruct-8da4w-gptq-8k.pte 8,192 tokens 0.80 GB passed ("Paris") with a completion prompt (the chat template did not render)
XNNPACK (CPU) any arm64 xnnpack/LFM2.5-1.2B-Instruct-8da4w-gptq-16k.pte 16,384 tokens 0.81 GB passed ("Paris") with a completion prompt (the chat template did not render)
XNNPACK (CPU) any arm64 xnnpack/LFM2.5-1.2B-Instruct-8da4w-gptq-32k.pte 32,768 tokens 0.83 GB passed ("Paris") with a completion prompt (the chat template did not render)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-512-chunk1of4.pte 512 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-512-chunk2of4.pte 512 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-512-chunk3of4.pte 512 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-512-chunk4of4.pte 512 tokens 0.41 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-2k-chunk1of4.pte 2,048 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-2k-chunk2of4.pte 2,048 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-2k-chunk3of4.pte 2,048 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-2k-chunk4of4.pte 2,048 tokens 0.41 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-4k-chunk1of4.pte 4,096 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-4k-chunk2of4.pte 4,096 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-4k-chunk3of4.pte 4,096 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-4k-chunk4of4.pte 4,096 tokens 0.41 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-8k-chunk1of4.pte 8,192 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-8k-chunk2of4.pte 8,192 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-8k-chunk3of4.pte 8,192 tokens 0.27 GB structure checked (no host NPU runtime)
MediaTek NeuroPilot MT6991 (Dimensity 9400) mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-a16w8-8k-chunk4of4.pte 8,192 tokens 0.41 GB structure checked (no host NPU runtime)

MediaTek folders also hold the token embedding table the NeuroPilot runner reads from disk, shared by every window: mtk/mt6991/LFM2.5-1.2B-Instruct-neuropilot-embedding-fp32.bin.

Tokenizer: tokenizer.json, copied unchanged from the source repo. Each backend folder has a config.json listing every window as a variant with the metadata the .pte reports, and an export-report-<window>.json per file with the full export record.

Memory

  • XNNPACK (CPU) at 2,048 tokens: the KV cache costs 24,576 bytes per token (fp32), 50,331,648 bytes for the whole window, allocated in full when the model loads.
  • XNNPACK (CPU) at 4,096 tokens: the KV cache costs 24,576 bytes per token (fp32), 100,663,296 bytes for the whole window, allocated in full when the model loads.
  • XNNPACK (CPU) at 8,192 tokens: the KV cache costs 24,576 bytes per token (fp32), 201,326,592 bytes for the whole window, allocated in full when the model loads.
  • XNNPACK (CPU) at 16,384 tokens: the KV cache costs 24,576 bytes per token (fp32), 402,653,184 bytes for the whole window, allocated in full when the model loads.
  • XNNPACK (CPU) at 32,768 tokens: the KV cache costs 24,576 bytes per token (fp32), 805,306,368 bytes for the whole window, allocated in full when the model loads.

How it was made

  • XNNPACK (CPU) any arm64: ExecuTorch 1.4.0 export_llm: 8-bit dynamic activations and 4-bit weights in groups of 32, int8 per-channel embeddings, XNNPACK with extended ops, prefill chunk 2048, fp32 KV cache. Built by run 1.
  • MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, lfm2.py): A16W8 (16-bit activations, 8-bit weights) calibrated on 23 samples in the model's own chat template filling the cache from 15 to 502 tokens (7 conversations from UltraChat 200k, under the app's system prompt where it fits, 7 passages from PG19 books, and MediaTek's 9 alpaca.txt prompts), 5,545 tokens observed, cut into 4 chunks, with a 128-token prompt graph and a one-token generation graph over a 512-token cache, compiled with MediaTek NeuroPilot Express SDK (mtk_converter 8.13.0+public, mtk_neuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.
  • MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, lfm2.py): A16W8 (16-bit activations, 8-bit weights) calibrated on 23 samples in the model's own chat template filling the cache from 15 to 2,038 tokens (7 conversations from UltraChat 200k, under the app's system prompt where it fits, 7 passages from PG19 books, and MediaTek's 9 alpaca.txt prompts), 17,365 tokens observed, cut into 4 chunks, with a 128-token prompt graph and a one-token generation graph over a 2048-token cache, compiled with MediaTek NeuroPilot Express SDK (mtk_converter 8.13.0+public, mtk_neuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.
  • MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, lfm2.py): A16W8 (16-bit activations, 8-bit weights) calibrated on 23 samples in the model's own chat template filling the cache from 15 to 4,085 tokens (7 conversations from UltraChat 200k, under the app's system prompt where it fits, 7 passages from PG19 books, and MediaTek's 9 alpaca.txt prompts), 32,724 tokens observed, cut into 4 chunks, with a 128-token prompt graph and a one-token generation graph over a 4096-token cache, compiled with MediaTek NeuroPilot Express SDK (mtk_converter 8.13.0+public, mtk_neuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.
  • MediaTek NeuroPilot MT6991 (Dimensity 9400): ExecuTorch 1.4.0 MediaTek LLM export (examples/mediatek, lfm2.py): A16W8 (16-bit activations, 8-bit weights) calibrated on 23 samples in the model's own chat template filling the cache from 15 to 8,182 tokens (7 conversations from UltraChat 200k, under the app's system prompt where it fits, 7 passages from PG19 books, and MediaTek's 9 alpaca.txt prompts), 63,149 tokens observed, cut into 4 chunks, with a 128-token prompt graph and a one-token generation graph over a 8192-token cache, compiled with MediaTek NeuroPilot Express SDK (mtk_converter 8.13.0+public, mtk_neuron 8.2.23) for MT6991 (Dimensity 9400). Built by run 1.

License

A quantized derivative of LiquidAI/LFM2.5-1.2B-Instruct, distributed under the same terms (lfm1.0). The upstream license files are included unchanged: LICENSE.

The mtk/ folders hold model binaries compiled with the MediaTek NeuroPilot Express SDK 8.0.8-build20250925 (mtk_converter 8.13.0+public, mtk_neuron 8.2.23) from MediaTek Inc., used under MediaTek's license terms for that SDK. No MediaTek SDK or runtime library is included. They run on MediaTek's LLM runner from ExecuTorch (examples/mediatek/executor_runner) with the device's NeuroPilot runtime; the settings it needs are in each folder's config.json, under each variant's runner.

Variants whose runner says state_layout: per-layer give each short-convolution layer its own two-position state, and take two small inputs that keep the runner's padding out of the convolution. MediaTek's runner as ExecuTorch 1.4.0 ships it cannot load them; the one that can is ExecuTorch's examples/mediatek/executor_runner/llama_runner with OpenWeights' patch, which tells caches from states by shape.

Downloads last month
438
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for experimentalmachines/LFM2.5-1.2B-Instruct-ExecuTorch

Quantized
(119)
this model

Collection including experimentalmachines/LFM2.5-1.2B-Instruct-ExecuTorch