Instructions to use SpacemiT/FastVLM-0.5B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use SpacemiT/FastVLM-0.5B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf SpacemiT/FastVLM-0.5B:Q4_1 # Run inference directly in the terminal: llama cli -hf SpacemiT/FastVLM-0.5B:Q4_1
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf SpacemiT/FastVLM-0.5B:Q4_1 # Run inference directly in the terminal: llama cli -hf SpacemiT/FastVLM-0.5B:Q4_1
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf SpacemiT/FastVLM-0.5B:Q4_1 # Run inference directly in the terminal: ./llama-cli -hf SpacemiT/FastVLM-0.5B:Q4_1
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf SpacemiT/FastVLM-0.5B:Q4_1 # Run inference directly in the terminal: ./build/bin/llama-cli -hf SpacemiT/FastVLM-0.5B:Q4_1
Use Docker
docker model run hf.co/SpacemiT/FastVLM-0.5B:Q4_1
- LM Studio
- Jan
- Ollama
How to use SpacemiT/FastVLM-0.5B with Ollama:
ollama run hf.co/SpacemiT/FastVLM-0.5B:Q4_1
- Unsloth Desktop
- Docker Model Runner
How to use SpacemiT/FastVLM-0.5B with Docker Model Runner:
docker model run hf.co/SpacemiT/FastVLM-0.5B:Q4_1
- Lemonade
How to use SpacemiT/FastVLM-0.5B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull SpacemiT/FastVLM-0.5B:Q4_1
Run and chat with the model
lemonade run user.FastVLM-0.5B-Q4_1
List all available models
lemonade list
- Atomic Chat
File size: 9,425 Bytes
7abc3c4 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 | ---
license: apple-amlr
pipeline_tag: image-to-text
library_name: llama.cpp
tags:
- fastvlm
- spacemit
- k1
- k3
- gguf
- onnxruntime
---
# FastVLM-0.5B for SpacemiT K1/K3
This repository contains a SpacemiT edge-deployment version of
[apple/FastVLM-0.5B](https://huggingface.co/apple/FastVLM-0.5B). The FastVLM
vision encoder is exported as an ONNX model and runs through the SpacemiT
Execution Provider, while the Qwen2-based 0.5B text decoder is stored as a
Q4_1 GGUF model and runs through the SpacemiT `llama.cpp` backend.
## Model origin and acknowledgement
FastVLM was introduced by Apple in
[FastVLM: Efficient Vision Encoding for Vision Language Models](https://arxiv.org/abs/2412.13303),
published at CVPR 2025. It is a vision-language model for understanding and
describing images. The key FastViTHD encoder is designed to produce fewer
visual tokens and reduce vision encoding latency, while retaining useful image
understanding quality. The FastVLM-0.5B variant is the compact model in the
released family.
This repository is a deployment conversion, not a new base model. We thank the
Apple Machine Learning Research team for releasing the FastVLM model, paper,
and source code:
- [FastVLM paper](https://arxiv.org/abs/2412.13303)
- [Official FastVLM repository](https://github.com/apple/ml-fastvlm)
- [Official FastVLM-0.5B model](https://huggingface.co/apple/FastVLM-0.5B)
The current SpacemiT package exposes offline image understanding through the
OpenAI-compatible `llama-server` chat completions API. The bundled
`humanspeech.jpg` sample has been validated on both K1 and K3.
## Files
```text
FastVLM-0.5B-SpacemiT/
├── fastvlm-text-0.5B-Q4_1.gguf
├── fastvlm_vision.f16.onnx
├── configs/
│ ├── K1/config.json
│ └── K3/config.json
├── humanspeech.jpg
└── README.md
```
The included `humanspeech.jpg` shows a woman speaking at a podium and is used
by the board smoke example below. `fastvlm_vision.f16.onnx` is the vision encoder;
`fastvlm-text-0.5B-Q4_1.gguf` is the text decoder.
## Supported platforms
| Platform | Accelerated cores | SMT config | `llama-server` threads |
|---|---:|---|---:|
| SpacemiT K1 / X60 | `0,1,2,3` | `configs/K1` | `-t 4` |
| SpacemiT K3 / A100 | `8,9,10,11,12,13,14,15` | `configs/K3` | `-t 8` |
The platform-specific `config.json` controls the ONNX vision encoder's
SpaceMIT EP thread count and affinity. The `-t` argument controls the GGUF
text decoder. Do not use the K3 config on K1: K1 has four accelerated cores at
IDs `0-3`, whereas K3 has eight accelerated cores at IDs `8-15`.
The SpacemiT `llama.cpp` runtime automatically detects the accelerated CPU
cores. Normal runs do not require manually setting
`SPACEMIT_PERFER_CORE_ARCH`, `SPACEMIT_PERFER_CORE_ID`, or
`SPACEMIT_MEM_BACKEND`.
## Prerequisites
Two runtime components are required:
1. A SpacemiT ONNX Runtime package containing `libonnxruntime` and the
SpacemiT Execution Provider:
[spacemit-com/onnxruntime releases](https://github.com/spacemit-com/onnxruntime/releases)
2. An SMT-enabled SpacemiT `llama.cpp` build containing `llama-server`:
[spacemit-com/llama.cpp](https://github.com/spacemit-com/llama.cpp)
The validation below used SpacemiT ORT `2.0.6` and a RISC-V
`llama-server` built from the SpacemiT fork.
### Option A: use prebuilt packages
Download and unpack the current RISC-V glibc releases:
```bash
wget https://github.com/spacemit-com/onnxruntime/releases/download/2.0.6/spacemit-ort.riscv64.2.0.6.tar.gz
tar -xf spacemit-ort.riscv64.2.0.6.tar.gz
wget https://github.com/spacemit-com/llama.cpp/releases/download/v0.1.7/spacemit-llama.cpp.riscv64.0.1.7.tar.gz
tar -xf spacemit-llama.cpp.riscv64.0.1.7.tar.gz
```
Check the release pages for newer mutually compatible packages when deploying
this model in another software image.
### Option B: build llama.cpp from source
Cross-compilation also requires a SpacemiT RISC-V toolchain:
```bash
git clone --recursive https://github.com/spacemit-com/llama.cpp.git
cd llama.cpp
export RISCV_ROOT_PATH=/path/to/spacemit-riscv-toolchain
export SPACEMIT_ORT_DIR=/path/to/spacemit-ort.riscv64.2.0.6
bash build_spacemit.sh glibc
```
The installed runtime is generated under `build/installed/`. An equivalent
manual CMake build must enable at least `GGML_CPU_RISCV64_SPACEMIT=ON` and
`LLAMA_SERVER_SMT_MTMD=ON`, and pass `SPACEMIT_ORT_DIR` to CMake.
## Run on a K1 or K3 board
Copy this model folder, the unpacked ORT package, and the prebuilt or locally
built `llama.cpp` installation to the board. Set the following paths to match
their actual locations:
```bash
export MODEL_DIR=/path/to/FastVLM-0.5B-SpacemiT
export ORT_DIR=/path/to/spacemit-ort.riscv64.2.0.6
export LLAMA_DIR=/path/to/spacemit-llama.cpp.riscv64.0.1.7
export LD_LIBRARY_PATH="${LLAMA_DIR}/lib:${ORT_DIR}/lib:${LD_LIBRARY_PATH:-}"
```
### K1
K1 uses four accelerated cores, `0-3`:
```bash
"${LLAMA_DIR}/bin/llama-server" \
-m "${MODEL_DIR}/fastvlm-text-0.5B-Q4_1.gguf" \
--media-backend smt \
--smt-config-dir "${MODEL_DIR}/configs/K1" \
-t 4 \
--host 0.0.0.0 \
--port 8080 \
--warmup
```
The K1 vision config applies:
```json
"ep_config": {
"SPACEMIT_EP_INTRA_THREAD_NUM": "4",
"SPACEMIT_EP_INTER_THREAD_NUM": "1",
"SPACEMIT_EP_INTRA_THREAD_AFFINITY": "0;1;2;3"
}
```
### K3
K3 uses eight accelerated cores, `8-15`:
```bash
"${LLAMA_DIR}/bin/llama-server" \
-m "${MODEL_DIR}/fastvlm-text-0.5B-Q4_1.gguf" \
--media-backend smt \
--smt-config-dir "${MODEL_DIR}/configs/K3" \
-t 8 \
--host 0.0.0.0 \
--port 8080 \
--warmup
```
The K3 vision config applies:
```json
"ep_config": {
"SPACEMIT_EP_INTRA_THREAD_NUM": "8",
"SPACEMIT_EP_INTER_THREAD_NUM": "1",
"SPACEMIT_EP_INTRA_THREAD_AFFINITY": "8;9;10;11;12;13;14;15"
}
```
Wait until the server prints a line similar to:
```text
llama_server: listening on http://0.0.0.0:8080
```
If the service is reachable outside a trusted local network, configure an API
key and suitable network access controls instead of exposing an unauthenticated
`0.0.0.0` endpoint.
## Send an image understanding request
The request can be sent from the board itself or another machine that can
reach the board. Set `SERVER_URL=http://127.0.0.1:8080` when testing locally,
or replace `BOARD_IP` when testing remotely.
```bash
export SERVER_URL=http://127.0.0.1:8080
base64 < "${MODEL_DIR}/humanspeech.jpg" | tr -d '\n' | jq -Rs '{
messages: [
{
role: "user",
content: [
{
type: "image_url",
image_url: {
url: ("data:image/jpeg;base64," + .)
}
},
{
type: "text",
text: "Describe the image content."
}
]
}
],
max_tokens: 64,
temperature: 0,
stream: false,
chat_template_kwargs: {
enable_thinking: false
}
}' | curl "${SERVER_URL}/v1/chat/completions" \
-H "Content-Type: application/json" \
--data-binary @-
```
The generated description is returned in `choices[0].message.content`. Replace
`humanspeech.jpg` with another JPEG/PNG image to describe your own image.
## Verified example
The bundled `humanspeech.jpg` was tested on 2026-08-10 with `--warmup` enabled.
Both boards returned HTTP 200 and generated a description of the speaker,
podium, microphone, clothing, and background banner. With `max_tokens=64`,
the responses were truncated at the requested token limit.
K1 example content:
```text
The image depicts a woman standing at a podium, delivering a speech or presentation. She is positioned at the center of the frame, with her right hand raised, holding a microphone close to her mouth, suggesting she is speaking. She is dressed in a patterned top and dark pants. Behind her, there is a banner
```
K3 example content:
```text
The image depicts a woman standing at a podium, delivering a speech or presentation. She is positioned at the center of the frame, with her right hand raised, holding a microphone close to her mouth, suggesting she is actively speaking. She is dressed in a patterned top and dark pants. Behind her, there is a
```
Observed single-request wall time after server warmup:
| Board | HTTP status | Wall time |
|---|---:|---:|
| K1 (`0-3`, `-t 4`) | 200 | 9.79 s |
| K3 (`8-15`, `-t 8`) | 200 | 1.91 s |
These are functional smoke-test observations rather than a formal benchmark.
Startup and first-request time can be longer while the GGUF model is loaded and
the ONNX graph is compiled.
## Citation
Please cite the original FastVLM work when using this converted model:
```bibtex
@InProceedings{fastvlm2025,
author = {Pavan Kumar Anasosalu Vasu and Fartash Faghri and Chun-Liang Li and Cem Koc and Nate True and Albert Antony and Gokul Santhanam and James Gabriel and Peter Grasch and Oncel Tuzel and Hadi Pouransari},
title = {FastVLM: Efficient Vision Encoding for Vision Language Models},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2025},
}
```
## License
The original FastVLM model is released under Apple's AMLR model license. See
the [official model card](https://huggingface.co/apple/FastVLM-0.5B) and the
official repository's `LICENSE_MODEL` for the terms that apply to the model.
The SpacemiT `llama.cpp` and ONNX Runtime packages are separate dependencies
and remain subject to their respective repository licenses.
|