Instructions to use cstr/LFM2.5-Encoder-230M-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use cstr/LFM2.5-Encoder-230M-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf cstr/LFM2.5-Encoder-230M-GGUF:F16 # Run inference directly in the terminal: llama cli -hf cstr/LFM2.5-Encoder-230M-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf cstr/LFM2.5-Encoder-230M-GGUF:F16 # Run inference directly in the terminal: llama cli -hf cstr/LFM2.5-Encoder-230M-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf cstr/LFM2.5-Encoder-230M-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf cstr/LFM2.5-Encoder-230M-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf cstr/LFM2.5-Encoder-230M-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf cstr/LFM2.5-Encoder-230M-GGUF:F16
Use Docker
docker model run hf.co/cstr/LFM2.5-Encoder-230M-GGUF:F16
- LM Studio
- Jan
- Ollama
How to use cstr/LFM2.5-Encoder-230M-GGUF with Ollama:
ollama run hf.co/cstr/LFM2.5-Encoder-230M-GGUF:F16
- Unsloth Desktop
- Docker Model Runner
How to use cstr/LFM2.5-Encoder-230M-GGUF with Docker Model Runner:
docker model run hf.co/cstr/LFM2.5-Encoder-230M-GGUF:F16
- Lemonade
How to use cstr/LFM2.5-Encoder-230M-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull cstr/LFM2.5-Encoder-230M-GGUF:F16
Run and chat with the model
lemonade run user.LFM2.5-Encoder-230M-GGUF-F16
List all available models
lemonade list
- Atomic Chat
LFM2.5-Encoder-230M โ calibrated CrispEmbed GGUFs
Quantizations of LiquidAI/LFM2.5-Encoder-230M for CrispEmbed. The upstream LFM Open License v1.0 is included unchanged as LICENSE. This is a bidirectional masked language model with 1024-dimensional token features, not a fine-tuned retrieval model. Multiple masks are evaluated jointly.
Files and measured reference agreement
CPU measurements against the original FP32 Python checkpoint, 15 probe cases:
| File | MB | Arithmetic | Minimum token cosine | Mask agreement |
|---|---|---|---|---|
| LFM2.5-Encoder-230M-crisp-Q4_K-imatrix.gguf | 165.3 | FP32 matrix casts | 0.889690 | 12/15 |
| LFM2.5-Encoder-230M-q4_ops_down_q8.gguf | 209.9 | FP32 row-dot | 0.928912 | 14/15 |
| LFM2.5-Encoder-230M-q8_ops_down_f16.gguf | 330.2 | FP32 row-dot | 0.999818 | 14/15 |
These counts measure reference agreement, not downstream accuracy. Profiles were
selected on these probes. The 210 MB precise-mode result is a numerical/MLM screen;
full API contracts were also checked in ordinary mode. The 330 MB precise mode
passes API contracts and numerical tolerances, but still changes one low-margin
masked prediction. Larger precision profiles were not consistently better.
Use the official F16 GGUF
for strongest checkpoint parity (0.999989 minimum token cosine, 15/15 masks).
The base alias lfm2-encoder-230m in CrispEmbed continues to select official F16.
Run
Build CrispEmbed from current main, including the published registry aliases (no release binary claim). Download and run the 330 MB model with:
CRISPEMBED_LFM2_F32_DOT=1 crispembed -m lfm2-encoder-230m-mixed330 \
--fill-mask "The capital of France is [MASK]." --json
The 210 MB alias is lfm2-encoder-230m-mixed210. The compact calibrated alias
is lfm2-encoder-230m-q4k; use CRISPEMBED_LFM2_F32_MATMUL=1 to reproduce its
reported FP32-cast measurement. Without these opt-ins the usual GGML matmuls
also quantize activation operands, and parity is lower. No GPU parity or
uncontended speed advantage is claimed. The precision modes change inference
arithmetic, not stored weights. CPU is the default for this encoder.
Raw token features are available through --raw-tokens, Python
encode_tokens(text, normalize=False), and the C/Rust/Dart APIs. Use raw features
for the tied MLM head; normalized or pooled features are not interchangeable.
The publisher's unchanged NumPy helper assumes floating-point weights; quantized
heads must be dequantized before projection.
Reproduce
Calibration and mixed precision commands, tensor override patterns, tokenizer
checks, decoded predictions, norms and limitations are documented in
CrispEmbed's encoder guide.
results/ contains the quantization, mixed-precision and arithmetic audit manifests.
The manifests include model SHA-256 hashes. Original FP32 and identical-Q8-weight
Python fixtures are hosted in
cstr/crispembed-regression-fixtures.
Both precise Q8 modes pass all 297 layer checks against identical-weight Python
(minimum token cosine 0.999999707, 15/15 same-weight masks).
The 165 MB profile used 6,711 calibration tokens. The mixed models used a fresh FP32 source and expanded independent calibration: 174 records / 185 samples / 12,065 tokens (maximum 1,149). Critical attention, ShortConv projection and FFN down matrices are kept at Q8_0 (210 MB) or F16 (330 MB); 28 FFN gate/up matrices remain calibrated Q4_K or Q8_0, respectively. The tied embedding/head is Q8_0; 49 F32 norm/kernel tensors are preserved. No training data or task fine-tuning was added. Calibration records and reproduction tooling live in CrispEmbed.
- Downloads last month
- -
16-bit
Model tree for cstr/LFM2.5-Encoder-230M-GGUF
Base model
LiquidAI/LFM2.5-230M-Base