LFM2-VL-450M โ€” Torq build (Synaptics SL2619 NPU)

Synaptics

This repository provides compiled model files for LiquidAI's LFM2-VL-450M vision-language model, ready to run on the Synaptics SL2610-series Torq NPU. Give it an image and a natural-language question, and it answers questions about that image.

Quick start guide:

  • Buy a Machina kit: Get an SL2600 Machina kit delivered to you
  • Torq Examples: Use Torq-examples LiquidAI/LiquidAI-LFM2-VL-450M scripts to download and deploy on your Machina kit

SL2600 Machina kit

Model Overview

LFM2โ€‘VL is designed to process text and images with variable resolutions. Built on the LFM2 backbone, it is optimized for low-latency and edge AI applications.

LFM2-VL utilizes hybrid conv/attention text decoders that execute on the NPU in bf16; the token embeddings run on the host CPU.

Image + prompt โ†’ caption / visual question answering. The image is encoded once and its KV cache is reused, so follow-up questions about the same image stay fast.

Model Features

Contents

File Size Size (W8) Role
vision_encoder_256.vmfb 203 MB 106 MB SigLIP vision encoder, 256-res โ†’ 64 image tokens
decoder_image_2part_A.vmfb 353 MB 150 MB one-shot image-prefill decoder, layers 0โ€“7
decoder_image_2part_B.vmfb 311 MB 132 MB one-shot image-prefill decoder, layers 8โ€“15
decoder_nolm.vmfb 577 MB 290 MB LFM2 single-token decode body (hidden-state output)
lm_head.vmfb 134 MB 67 MB tied LM head (hidden โ†’ 65 536 logits)
token_embeddings.npy 134 MB โ€” CPU embedding LUT / tied-LM-head weights (bf16)
config.json, tokenizer.json โ€” โ€” model config + tokenizer
cats-and-dogs-256.jpg โ€” โ€” sample 256-res image for the demo
onnx/ ~2 GB โ€” reference ONNX exports (vision encoder, merged decoder, embeddings) for non-Torq runtimes

Model Details

  • Base model: LiquidAI LFM2-VL-450M (SigLIP vision tower + LFM2 language model).
  • Text decoder: LFM2 โ€” hidden size 1024, 16 layers, 16 attention heads, vocabulary 65 536, hybrid short-convolution + grouped-query attention.
  • Image tokens: 64 per image (256-resolution input).
  • Precision: bf16 on the NPU.
  • Target: Synaptics SL2619, compiled with the Torq compiler.
  • On-device performance (SL2619, indicative): vision encode ~2.4 s, imageโ†’KV prefill ~3.7 s, decode ~3.6โ€“4.2 tok/s.

Tested Platforms

Metrics

Platform Model / Stage Environment NPU Clock TTFT Infer / s
SL2619 2GB LFM2-VL-450M Torq v2.0.0 1 GHz 2844 ms 3.4
SL2619 2GB LFM2-VL-450M-W8 Torq v2.0.0 1 GHz 1371 ms 6.2

Deployment

The models have been tested with the following environment.

  • Torq Compiler: v2.0.0
  • Torq Runtime: v2.0.0 included in Astra SDK release scarthgap_6.12_v2.4.0

Usage Tutorials / Example Apps

A usage example is provided in the Torq Examples / LiquidAI-LFM2-VL-450M.

Check out the README for instructions.

License & attribution

This repository is a redistribution of a model created by Liquid AI, Inc., licensed under the LFM Open License v1.0. Copies of the license and the attribution notices are included alongside the model files:

  • LICENSE โ€” a verbatim copy of the LFM Open License v1.0.
  • NOTICE โ€” the copyright, patent, trademark, and attribution notices retained from the original Work (per Section 4(c) of the license).

Original model: LFM2.5-230M ยท Copyright ยฉ Liquid AI, Inc.

Learn More

Downloads last month
109
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Synaptics/LiquidAI-LFM2-VL-450M

Quantized
(18)
this model