Instructions to use itsZyn/ZynDwarf-1.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use itsZyn/ZynDwarf-1.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="itsZyn/ZynDwarf-1.1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("itsZyn/ZynDwarf-1.1") model = AutoModelForCausalLM.from_pretrained("itsZyn/ZynDwarf-1.1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use itsZyn/ZynDwarf-1.1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf itsZyn/ZynDwarf-1.1:Q4_K_M # Run inference directly in the terminal: llama cli -hf itsZyn/ZynDwarf-1.1:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf itsZyn/ZynDwarf-1.1:Q4_K_M # Run inference directly in the terminal: llama cli -hf itsZyn/ZynDwarf-1.1:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf itsZyn/ZynDwarf-1.1:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf itsZyn/ZynDwarf-1.1:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf itsZyn/ZynDwarf-1.1:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf itsZyn/ZynDwarf-1.1:Q4_K_M
Use Docker
docker model run hf.co/itsZyn/ZynDwarf-1.1:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use itsZyn/ZynDwarf-1.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "itsZyn/ZynDwarf-1.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "itsZyn/ZynDwarf-1.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/itsZyn/ZynDwarf-1.1:Q4_K_M
- SGLang
How to use itsZyn/ZynDwarf-1.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "itsZyn/ZynDwarf-1.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "itsZyn/ZynDwarf-1.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "itsZyn/ZynDwarf-1.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "itsZyn/ZynDwarf-1.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use itsZyn/ZynDwarf-1.1 with Ollama:
ollama run hf.co/itsZyn/ZynDwarf-1.1:Q4_K_M
- Unsloth Desktop
- Pi
How to use itsZyn/ZynDwarf-1.1 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf itsZyn/ZynDwarf-1.1:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "itsZyn/ZynDwarf-1.1:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use itsZyn/ZynDwarf-1.1 with Docker Model Runner:
docker model run hf.co/itsZyn/ZynDwarf-1.1:Q4_K_M
- Lemonade
How to use itsZyn/ZynDwarf-1.1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull itsZyn/ZynDwarf-1.1:Q4_K_M
Run and chat with the model
lemonade run user.ZynDwarf-1.1-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use itsZyn/ZynDwarf-1.1 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf itsZyn/ZynDwarf-1.1:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default itsZyn/ZynDwarf-1.1:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use itsZyn/ZynDwarf-1.1 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf itsZyn/ZynDwarf-1.1:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "itsZyn/ZynDwarf-1.1:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- ZynDwarf-1.1
- What is ZynDwarf-1.1?
- Release files
- Model specification
- Agent and tool calling
- Ollama agent compatibility
- Real local measurements
- Published comparison with other small models
- Broader upstream comparison
- Training details
- Context length
- Transformers
- llama.cpp
- Ollama
- OpenAI-compatible local APIs
- Using ZynDwarf in an agent loop
- Security model
- Evaluation methodology
- Limitations
- Known failure modes
- Why F16 and Q4 are both released
- Reproducibility
- Version history
- Credits and references
- Citation
- Final release statement
ZynDwarf-1.1
ZynDwarf-1.1 is a compact, text-only general assistant built for local use, programming, practical reasoning, and agent-oriented workflows.
It is derived from Liquid AI's LFM2.5-350M family and adapted by Zyn Models / itsZyn. The release keeps the small footprint of the base model while adding a training focus on everyday conversation, Spanish and English interaction, coding, debugging, structured output, tool selection, tool-call formatting, tool-result continuation, recovery from tool failures, and multi-turn agent behavior.
The point of this release is not to pretend that a 354M-parameter model is secretly a 70B model wearing a tiny coat. The point is to make a genuinely useful small model that can be deployed on CPU-oriented hardware, embedded into local tools, and connected to real agent runtimes without requiring a large inference stack.
Release: 1.1
Model family: LFM2 / LFM2.5
Parameters: ~354.5M
Primary language focus: Spanish + English
Primary local formats: Safetensors, GGUF F16, GGUF Q4_K_M
Agent runtime targets: Transformers, llama.cpp, Ollama
Table of contents
- What is ZynDwarf-1.1?
- Release files
- Model specification
- Agent and tool calling
- Ollama agent compatibility
- Real local measurements
- Published comparison with other small models
- Training details
- Context length
- Transformers
- llama.cpp
- Ollama
- OpenAI-compatible local APIs
- Using ZynDwarf in an agent loop
- Security model
- Evaluation methodology
- Limitations
- Known failure modes
- Why F16 and Q4 are both released
- Reproducibility
- Version history
- Credits and references
- Citation
What is ZynDwarf-1.1?
ZynDwarf-1.1 is a small language model release intended to sit between a plain chat model and a purpose-built autonomous agent.
Its most useful design assumption is simple:
the model should decide when a tool is useful, emit a structured call when a tool is actually needed, and leave execution to the host application.
This matters because an agent is not merely a model with the word "agent" pasted into the README. A reliable agent is a loop consisting of a model, tool definitions, validation, execution, observations, and another model turn.
ZynDwarf-1.1 is therefore packaged to support that loop without giving the model unrestricted access to the machine running it.
Typical workloads include:
- everyday conversation;
- explanations in Spanish or English;
- Python, JavaScript, TypeScript, Bash, HTML/CSS, JSON, SQL and configuration work;
- debugging and code transformation;
- lightweight planning;
- filesystem or shell tool selection through a host runtime;
- structured responses;
- local CPU inference;
- small autonomous or semi-autonomous agents where the surrounding system performs the important validation.
The model is intentionally not marketed as a large reasoning specialist, a research model, or a replacement for substantially larger systems.
Release files
The repository contains the full Transformers release and both primary GGUF variants.
| File | Format | Approx. size | Intended use |
|---|---|---|---|
model-00001-of-00003.safetensors |
Safetensors | ~475 MB | Transformers |
model-00002-of-00003.safetensors |
Safetensors | ~471 MB | Transformers |
model-00003-of-00003.safetensors |
Safetensors | ~409 MB | Transformers |
ZynDwarf-1.1-f16.gguf |
F16 GGUF | 676.25 MiB | Reference-quality local inference |
ZynDwarf-1.1-Q4_K_M.gguf |
Q4_K_M GGUF | 216.41 MiB | Lower-memory local inference |
chat_template.jinja |
Jinja template | small | Chat + tool rendering |
agent_config.json |
JSON | small | Agent runtime metadata |
SHA256SUMS |
text | small | Artifact verification |
The two GGUF files are deliberately distributed next to the Transformers release so a Hugging Face user can discover both precision options from one model page.
The Q4_K_M artifact in this release was quantized directly from the released F16 artifact with llama.cpp's llama-quantize using Q4_K_M.
Model specification
| Property | ZynDwarf-1.1 |
|---|---|
| Base model | LiquidAI/LFM2.5-350M |
| Model type | lfm2 |
| Architecture | Lfm2ForCausalLM |
| Parameter count | ~354.5M |
| Vocabulary | 65,536 |
| Architectural maximum position length | 128,000 tokens |
| Training sequence cap | 768 tokens |
| Recommended runtime context | 32,768 tokens |
| LoRA rank | 8 |
| LoRA alpha | 16 |
| LoRA dropout | 0.05 |
| LoRA target modules | q / k / v projections |
| Learning rate | 1.5e-6 |
| Epochs | 1 |
| Gradient accumulation | 4 |
| Optimizer steps | 376 |
| Encoded training sequences | 1,503 |
| Tool-oriented examples | 308 |
| Multi-turn examples | 62 |
| Runtime targets | Transformers / llama.cpp / Ollama |
| Vision | No |
| Audio | No |
| Text-to-image | No |
| Native tool-call template | Yes |
Agent and tool calling
Native format
The model's chat template knows how to render tool definitions and assistant tool calls.
The native assistant form is:
<|tool_call_start|>[ToolName(arg=value)]<|tool_call_end|>
For example:
<|tool_call_start|>[Bash(command='free -h')]<|tool_call_end|>
The exact function-call object presented to an application can differ by runtime. Transformers exposes model-native content through the chat template, while Ollama parses its native tool-call response into message.tool_calls objects.
Agent loop
A normal host-controlled loop looks like this:
┌──────────────┐
│ User request │
└──────┬───────┘
│
v
┌──────────────────────┐
│ Model + tool schemas │
└──────────┬───────────┘
│
v
┌─────────────────┐
│ Final answer? │────── yes ──────> return answer
└────────┬────────┘
│ no
v
┌─────────────────┐
│ Tool call │
└────────┬────────┘
│
v
┌─────────────────┐
│ Host validation │
└────────┬────────┘
│
v
┌─────────────────┐
│ Execute tool │
└────────┬────────┘
│
v
┌─────────────────┐
│ Tool result │
└────────┬────────┘
│
└──────────────> next model turn
The host application remains the security boundary.
The model generates requests. The host decides whether those requests are allowed, what they mean, whether the arguments are valid, and what execution result is returned.
What the model does not do
ZynDwarf-1.1 does not receive direct permission to execute Bash, access the filesystem, modify databases, send network requests, or change the host operating system.
Those capabilities come from tools supplied by the host.
That distinction is important for every agent integration, particularly on machines where a mistaken tool call could have real consequences.
Ollama agent compatibility
Ollama supports tool calling through its /api/chat endpoint and accepts tool schemas alongside messages. See the official Ollama tool-calling documentation:
https://docs.ollama.com/capabilities/tool-calling
ZynDwarf-1.1 was tested locally through this interface with the same host-controlled pattern.
Real smoke test
Tool supplied by the host:
{
"type": "function",
"function": {
"name": "free_memory",
"description": "Return the current available RAM in the system",
"parameters": {
"type": "object",
"properties": {},
"required": []
}
}
}
The model returned a structured call whose function name was free_memory for both published variants.
Measured local result
| Variant | Native tool-call smoke | Observed |
|---|---|---|
| F16 | 1 / 1 | message.tool_calls[0].function.name = free_memory |
| Q4_K_M | 1 / 1 | message.tool_calls[0].function.name = free_memory |
This is a smoke test, not a claim that the model will correctly call every arbitrary tool on every prompt. The purpose is to verify that the released GGUF + Ollama packaging participates correctly in a real structured tool-call path.
Real local measurements
CPU throughput
The release was measured using llama-bench from llama.cpp on the development server.
Test conditions:
- CPU: 2-vCPU Intel Xeon host;
- threads: 2;
- prompt tokens: 128;
- generated tokens: 128;
- llama.cpp build:
b19cbe9, build 8; - measurement type: prompt processing (
pp128) and token generation (tg128).
Raw result
| Variant | File size | Prompt processing | Generation |
|---|---|---|---|
| F16 | 676.25 MiB | 221.00 ± 7.96 tok/s | 16.51 ± 2.13 tok/s |
| Q4_K_M | 216.41 MiB | 281.80 ± 16.07 tok/s | 38.52 ± 0.80 tok/s |
Q4_K_M is approximately 68% smaller than F16 on disk and was materially faster in this CPU benchmark.
The throughput figures are machine-specific measurements. They are not a standardized model-quality benchmark and must not be compared to benchmark numbers quoted by another model vendor under a different GPU, CPU, context length, quantization, batch size, tokenizer, or runtime.
Q4 compression ratio
Using the measured GGUF sizes:
- F16: 676.25 MiB
- Q4_K_M: 216.41 MiB
- reduction: about 68%
This is why Q4_K_M is the recommended default for smaller local machines, while F16 remains the reference artifact for fidelity comparisons.
Published comparison with other small models
The comparison below is intentionally split into two groups:
- our real local measurements, such as throughput and the Ollama tool-call smoke test;
- public benchmark numbers published by the upstream model authors.
These are not a single leaderboard.
A local 2-thread CPU throughput test and a benchmark score reported from an H100 cluster are measuring completely different things.
IFEval snapshot
IFEval is useful here because it focuses on instruction following rather than raw language-model memorization.
Liquid AI's current model card reports:
| Model | Approx. parameters | IFEval |
|---|---|---|
| LFM2.5-350M | 0.35B | 76.96 |
| Granite 4.0-H-350M | 0.35B | 61.27 |
| Granite 4.0-350M | 0.35B | 53.48 |
Source:
https://huggingface.co/LiquidAI/LFM2.5-350M
The official SmolLM2 model card reports this compact-model comparison under its lighteval setup:
| Model | Approx. parameters | IFEval |
|---|---|---|
| SmolLM2-360M-Instruct | 0.36B | 41.0 |
| Qwen2.5-0.5B-Instruct | 0.50B | 31.6 |
Source:
https://huggingface.co/HuggingFaceTB/SmolLM2-360M-Instruct
Why ZynDwarf-1.1 is not assigned an IFEval number here
We did not run the IFEval benchmark itself in this release environment.
Instead, we ran real local runtime tests and smoke evaluations.
Assigning a made-up IFEval percentage by converting a different 1-to-1 smoke test into a leaderboard score would not be an evaluation. It would be decorative arithmetic.
The honest release therefore keeps the published upstream numbers separate from our own measurements.
Broader upstream comparison
Liquid AI's published model table is particularly useful because it places LFM2.5-350M directly beside models of a similar footprint.
| Model | GPQA Diamond | MMLU-Pro | IFEval | BFCLv3 | BFCLv4 |
|---|---|---|---|---|---|
| LFM2.5-350M | 30.64 | 20.01 | 76.96 | 44.11 | 21.86 |
| LFM2-350M | 27.58 | 19.29 | 64.96 | 22.95 | 12.29 |
| Granite 4.0-H-350M | 22.32 | 13.14 | 61.27 | 43.07 | 13.28 |
| Granite 4.0-350M | 25.91 | 12.84 | 53.48 | 39.58 | 13.73 |
| Gemma 3 1B IT | 23.89 | 14.04 | 63.49 | 16.61 | 7.17 |
Source:
https://huggingface.co/LiquidAI/LFM2.5-350M
These numbers should be interpreted as evidence about the upstream model family and its published evaluation, not as measured ZynDwarf-1.1 scores.
ZynDwarf-1.1 inherits the same underlying LFM2.5-350M architecture and tokenizer family, but it has been adapted with a separate behavioral training run.
Training details
Starting point
Base model:
LiquidAI/LFM2.5-350M
ZynDwarf-1.1 was not trained from scratch.
The release started from a repaired LFM2.5-350M-derived checkpoint and then applied a general agent-oriented LoRA adaptation.
Dataset
The intended training dataset contained:
- 1,613 deduplicated sequences;
- 1,503 sequences that successfully entered the encoder under the 768-token cap;
- 308 examples containing real tool-oriented traces;
- 62 multi-turn examples.
The behavioral mix covered:
- normal conversation;
- Spanish responses;
- English responses;
- programming;
- debugging;
- shell and Linux tasks;
- structured output;
- tool selection;
- tool-call formatting;
- multi-step work;
- tool failure recovery;
- cases where no tool should be used;
- planning and task decomposition.
LoRA configuration
rank = 8
alpha = 16
dropout = 0.05
targets = q_proj, k_proj, v_proj
learning_rate = 1.5e-6
epochs = 1
grad_accum = 4
max_length = 768
The adapter trained approximately 0.069% of the total model parameter count.
What this release does not claim
This release does not claim:
- logit-level knowledge distillation;
- 128K-token end-to-end training;
- perfect tool use;
- perfect reasoning;
- broad factual reliability;
- autonomous execution without a host;
- parity with larger coding specialists.
The documentation is intentionally explicit about those boundaries.
Context length
There are three different concepts that are easy to accidentally collapse into one number.
1. Architectural maximum
The released Transformers configuration exposes a maximum position length of 128,000 tokens.
2. Training sequence cap
The fine-tuning run used a maximum encoded sequence length of 768 tokens.
3. Recommended runtime context
For practical local inference, the published examples use 32,768 tokens.
These values describe different parts of the system.
A model having 128K positional metadata does not mean that it was fine-tuned end-to-end on 128K-token conversations, nor does it mean that a 3.8 GiB RAM machine can comfortably process 128K tokens.
The release therefore treats 32K as the conservative runtime setting and 128K as an architectural configuration value.
Transformers
Installation
pip install -U transformers torch
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "itsZyn/ZynDwarf-1.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
messages = [
{"role": "user", "content": "Hola, ¿qué puedes hacer?"}
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
)
outputs = model.generate(
inputs,
max_new_tokens=128,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=False))
Tool schemas
When the runtime and Transformers version support tool-aware chat templates, pass tool definitions through the tools= argument.
messages = [
{"role": "user", "content": "Check the current memory."}
]
inputs = tokenizer.apply_chat_template(
messages,
tools=tools,
add_generation_prompt=True,
return_tensors="pt",
)
The host should parse and validate any generated tool call before execution.
llama.cpp
The GGUF files can be used directly with a current llama.cpp build that supports the LFM2 architecture.
F16
./llama-cli \
-m ZynDwarf-1.1-f16.gguf \
-c 32768 \
-t 2 \
--temp 0.35 \
--top-k 40 \
--top-p 0.9 \
--repeat-penalty 1.05
Q4_K_M
./llama-cli \
-m ZynDwarf-1.1-Q4_K_M.gguf \
-c 32768 \
-t 2 \
--temp 0.35 \
--top-k 40 \
--top-p 0.9 \
--repeat-penalty 1.05
For production agent work, prefer the runtime's structured chat API rather than manually concatenating raw control tokens unless you know exactly how that runtime handles the template.
Ollama
Pull the default Q4 release
ollama pull itsZyn/ZynDwarf-1.1:latest
Run
ollama run itsZyn/ZynDwarf-1.1:latest
F16 tag
ollama pull itsZyn/ZynDwarf-1.1:f16
Q4 tag
ollama pull itsZyn/ZynDwarf-1.1:q4_k_m
The default latest tag is the Q4_K_M build because that is the most practical variant for small local machines.
The f16 tag is the higher-fidelity reference build.
The q4_k_m tag is the compact build.
OpenAI-compatible local APIs
Ollama exposes an OpenAI-compatible endpoint. The model can therefore be used by applications that speak the OpenAI chat-completions protocol, subject to the specific client's tool-call support.
Example:
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "itsZyn/ZynDwarf-1.1:latest",
"messages": [
{"role": "user", "content": "Hello"}
]
}'
For structured agent use, the native Ollama /api/chat endpoint is preferable because it exposes the tools field directly:
https://docs.ollama.com/api/chat
Using ZynDwarf in an agent loop
A minimal conceptual Python loop looks like this:
from ollama import chat
messages = [
{"role": "user", "content": "Check the current memory."}
]
def free_memory():
# host-side implementation
...
tools = [free_memory]
while True:
response = chat(
model="itsZyn/ZynDwarf-1.1:latest",
messages=messages,
tools=tools,
)
messages.append(response.message)
if not response.message.tool_calls:
print(response.message.content)
break
for call in response.message.tool_calls:
# validate the function name and arguments before execution
if call.function.name == "free_memory":
result = free_memory()
else:
raise ValueError("Unknown tool")
messages.append({
"role": "tool",
"tool_name": call.function.name,
"content": str(result),
})
The example is intentionally host-centric. The model requests an operation; the application decides whether to execute it.
Security model
Treat every model-generated tool call as untrusted input.
At minimum, an agent runtime should validate:
- tool name;
- argument names;
- argument types;
- path restrictions;
- allowed commands;
- network destinations;
- maximum execution time;
- maximum output size;
- permissions;
- whether the requested action is actually necessary.
The model should not be the component that decides its own authorization policy.
For filesystem tools, prefer allowlisted project directories over unrestricted root access.
For shell tools, prefer narrowly defined functions over an unrestricted command interpreter.
For network tools, restrict destinations where possible.
For destructive operations, require an explicit policy or confirmation outside the model.
These practices are properties of the agent runtime rather than special powers granted by ZynDwarf itself.
Evaluation methodology
The release follows three separate evidence levels.
Level A: artifact validation
The GGUF files were loaded by llama.cpp and benchmarked successfully.
This verifies that the artifacts are structurally usable by the target runtime.
Level B: local runtime smoke tests
The F16 and Q4 Ollama packages were exercised through the actual /api/chat endpoint with a real tool schema.
Both returned a structured tool call for the test function.
This verifies the packaging path:
host tool schema -> Ollama -> ZynDwarf-1.1 -> structured tool call
Level C: upstream standardized benchmarks
Where standardized benchmark numbers are cited, they are labeled as upstream published results and linked to the originating model card.
No local smoke score is silently converted into IFEval, BFCL, MMLU, GPQA, GSM8K, HumanEval, or another standardized benchmark.
Limitations
ZynDwarf-1.1 is a small model.
At roughly 354.5M parameters, it has less representational capacity than much larger models. That affects:
- difficult mathematics;
- long-horizon planning;
- broad factual recall;
- deep debugging;
- large repositories;
- complicated multi-step reasoning;
- unusual tool-use schemas;
- long conversations with many state changes.
It can also make plausible mistakes.
Generated code still requires review.
Generated tool calls still require validation.
The model has no independent access to current information unless a host supplies a tool or retrieval system.
It is not intended for medical, legal, financial, security-critical, or other high-stakes decision making.
Known failure modes
A compact model can fail in ways that look surprisingly confident.
Common risk classes include:
Arithmetic drift
Short arithmetic problems can still produce incorrect intermediate reasoning, especially when the prompt requests an explanation instead of a single result.
Strict JSON failure
The model can occasionally add Markdown fences or explanatory text when the user requested JSON-only output. For applications requiring machine-readable output, validate the response and consider a structured-output-capable runtime.
Tool over- or under-use
The model may sometimes answer conceptually when a tool should have been used, or propose a tool when a direct answer would be sufficient. Agent runtimes should therefore implement a tool policy outside the model as well.
Context degradation
A large configured context window does not guarantee stable quality throughout the full window.
Quantization differences
Q4_K_M is faster and smaller, but it is not numerically identical to F16.
The F16 artifact should be treated as the reference model for fidelity investigations.
Why F16 and Q4 are both released
The two variants solve different deployment problems.
F16
Choose F16 when:
- memory is available;
- output fidelity matters more than disk size;
- you are comparing future model changes;
- you want a reference GGUF.
Q4_K_M
Choose Q4_K_M when:
- RAM is constrained;
- CPU inference matters;
- local storage is limited;
- you are building a lightweight agent on a modest device.
The current benchmark shows the expected trade-off clearly:
- F16: 676.25 MiB, 16.51 tok/s generation;
- Q4_K_M: 216.41 MiB, 38.52 tok/s generation;
on the exact test host and exact llama.cpp build documented above.
Reproducibility
To reproduce the release packaging, keep the following information together:
Base model: LiquidAI/LFM2.5-350M
LoRA rank: 8
LoRA alpha: 16
LoRA dropout: 0.05
Targets: q_proj, k_proj, v_proj
Learning rate: 1.5e-6
Epochs: 1
Gradient accumulation: 4
Max training sequence: 768
Encoded sequences: 1,503
Tool-oriented examples: 308
Multi-turn examples: 62
GGUF converter/runtime benchmark build: b19cbe9 (build 8)
F16 SHA256: ce452090b010d1e199fd7b39b2d65e5fcd8d7c2ceac52c38219f8b53e491c424
Q4 SHA256: 04a6ea0100b84a8687162db856f4c0b8042f54ad6b50a93b3c8132afc3499ae2
Always verify downloaded artifacts against SHA256SUMS before using them in a reproducible pipeline.
Version history
ZynDwarf-1.1
This release supersedes the public General Agent publication.
Highlights:
- unified model name;
- unified Hugging Face repository;
- unified Ollama repository;
- F16 + Q4_K_M GGUF distribution;
- explicit agent metadata;
- native tool-call documentation;
- real Ollama structured tool-call smoke test;
- refreshed llama.cpp throughput measurements;
- published comparison against compact model peers;
- reproducibility information;
- explicit limitations and benchmark separation.
General Agent publication
The older General Agent naming was retired in favor of the versioned ZynDwarf-1.1 release line.
Credits and references
Base model
Liquid AI, LFM2.5-350M:
https://huggingface.co/LiquidAI/LFM2.5-350M
The upstream model card documents the base model architecture, supported languages, context configuration, tool-use format, and published evaluation results.
SmolLM2 comparison
Hugging Face, SmolLM2-360M-Instruct:
https://huggingface.co/HuggingFaceTB/SmolLM2-360M-Instruct
The model card provides published compact-model comparisons including SmolLM2-360M-Instruct and Qwen2.5-0.5B-Instruct.
Ollama tool calling
Official Ollama tool-calling documentation:
https://docs.ollama.com/capabilities/tool-calling
Ollama API
Official Ollama chat API:
https://docs.ollama.com/api/chat
Ollama model format
Official Ollama model import documentation:
https://docs.ollama.com/import
llama.cpp
https://github.com/ggml-org/llama.cpp
Citation
@misc{zyndwarf11,
title = {ZynDwarf-1.1},
author = {Zyn Models and itsZyn},
year = {2026},
note = {Compact LFM2.5-350M based conversational and agent-oriented model}
}
Final release statement
ZynDwarf-1.1 is a compact local-first model for practical assistant and agent workflows.
It is intentionally small, intentionally measurable, and intentionally explicit about what has and has not been tested.
The F16 and Q4_K_M artifacts are released together.
The Ollama package is tested with real structured tool calls.
The throughput numbers are measured on the actual development CPU.
The external benchmark numbers are labeled as upstream numbers instead of being presented as if they were produced locally.
That is the release philosophy: small model, real runtime, real files, real measurements, and no imaginary leaderboard points.
- Downloads last month
- 1,258