ZynDwarf-1.1

ZynDwarf-1.1 is a compact, text-only general assistant built for local use, programming, practical reasoning, and agent-oriented workflows.

It is derived from Liquid AI's LFM2.5-350M family and adapted by Zyn Models / itsZyn. The release keeps the small footprint of the base model while adding a training focus on everyday conversation, Spanish and English interaction, coding, debugging, structured output, tool selection, tool-call formatting, tool-result continuation, recovery from tool failures, and multi-turn agent behavior.

The point of this release is not to pretend that a 354M-parameter model is secretly a 70B model wearing a tiny coat. The point is to make a genuinely useful small model that can be deployed on CPU-oriented hardware, embedded into local tools, and connected to real agent runtimes without requiring a large inference stack.

Release: 1.1
Model family: LFM2 / LFM2.5
Parameters: ~354.5M
Primary language focus: Spanish + English
Primary local formats: Safetensors, GGUF F16, GGUF Q4_K_M
Agent runtime targets: Transformers, llama.cpp, Ollama


Table of contents

  1. What is ZynDwarf-1.1?
  2. Release files
  3. Model specification
  4. Agent and tool calling
  5. Ollama agent compatibility
  6. Real local measurements
  7. Published comparison with other small models
  8. Training details
  9. Context length
  10. Transformers
  11. llama.cpp
  12. Ollama
  13. OpenAI-compatible local APIs
  14. Using ZynDwarf in an agent loop
  15. Security model
  16. Evaluation methodology
  17. Limitations
  18. Known failure modes
  19. Why F16 and Q4 are both released
  20. Reproducibility
  21. Version history
  22. Credits and references
  23. Citation

What is ZynDwarf-1.1?

ZynDwarf-1.1 is a small language model release intended to sit between a plain chat model and a purpose-built autonomous agent.

Its most useful design assumption is simple:

the model should decide when a tool is useful, emit a structured call when a tool is actually needed, and leave execution to the host application.

This matters because an agent is not merely a model with the word "agent" pasted into the README. A reliable agent is a loop consisting of a model, tool definitions, validation, execution, observations, and another model turn.

ZynDwarf-1.1 is therefore packaged to support that loop without giving the model unrestricted access to the machine running it.

Typical workloads include:

  • everyday conversation;
  • explanations in Spanish or English;
  • Python, JavaScript, TypeScript, Bash, HTML/CSS, JSON, SQL and configuration work;
  • debugging and code transformation;
  • lightweight planning;
  • filesystem or shell tool selection through a host runtime;
  • structured responses;
  • local CPU inference;
  • small autonomous or semi-autonomous agents where the surrounding system performs the important validation.

The model is intentionally not marketed as a large reasoning specialist, a research model, or a replacement for substantially larger systems.


Release files

The repository contains the full Transformers release and both primary GGUF variants.

File Format Approx. size Intended use
model-00001-of-00003.safetensors Safetensors ~475 MB Transformers
model-00002-of-00003.safetensors Safetensors ~471 MB Transformers
model-00003-of-00003.safetensors Safetensors ~409 MB Transformers
ZynDwarf-1.1-f16.gguf F16 GGUF 676.25 MiB Reference-quality local inference
ZynDwarf-1.1-Q4_K_M.gguf Q4_K_M GGUF 216.41 MiB Lower-memory local inference
chat_template.jinja Jinja template small Chat + tool rendering
agent_config.json JSON small Agent runtime metadata
SHA256SUMS text small Artifact verification

The two GGUF files are deliberately distributed next to the Transformers release so a Hugging Face user can discover both precision options from one model page.

The Q4_K_M artifact in this release was quantized directly from the released F16 artifact with llama.cpp's llama-quantize using Q4_K_M.


Model specification

Property ZynDwarf-1.1
Base model LiquidAI/LFM2.5-350M
Model type lfm2
Architecture Lfm2ForCausalLM
Parameter count ~354.5M
Vocabulary 65,536
Architectural maximum position length 128,000 tokens
Training sequence cap 768 tokens
Recommended runtime context 32,768 tokens
LoRA rank 8
LoRA alpha 16
LoRA dropout 0.05
LoRA target modules q / k / v projections
Learning rate 1.5e-6
Epochs 1
Gradient accumulation 4
Optimizer steps 376
Encoded training sequences 1,503
Tool-oriented examples 308
Multi-turn examples 62
Runtime targets Transformers / llama.cpp / Ollama
Vision No
Audio No
Text-to-image No
Native tool-call template Yes

Agent and tool calling

Native format

The model's chat template knows how to render tool definitions and assistant tool calls.

The native assistant form is:

<|tool_call_start|>[ToolName(arg=value)]<|tool_call_end|>

For example:

<|tool_call_start|>[Bash(command='free -h')]<|tool_call_end|>

The exact function-call object presented to an application can differ by runtime. Transformers exposes model-native content through the chat template, while Ollama parses its native tool-call response into message.tool_calls objects.

Agent loop

A normal host-controlled loop looks like this:

┌──────────────┐
│ User request │
└──────┬───────┘
       │
       v
┌──────────────────────┐
│ Model + tool schemas  │
└──────────┬───────────┘
           │
           v
  ┌─────────────────┐
  │ Final answer?   │────── yes ──────> return answer
  └────────┬────────┘
           │ no
           v
  ┌─────────────────┐
  │ Tool call       │
  └────────┬────────┘
           │
           v
  ┌─────────────────┐
  │ Host validation │
  └────────┬────────┘
           │
           v
  ┌─────────────────┐
  │ Execute tool    │
  └────────┬────────┘
           │
           v
  ┌─────────────────┐
  │ Tool result     │
  └────────┬────────┘
           │
           └──────────────> next model turn

The host application remains the security boundary.

The model generates requests. The host decides whether those requests are allowed, what they mean, whether the arguments are valid, and what execution result is returned.

What the model does not do

ZynDwarf-1.1 does not receive direct permission to execute Bash, access the filesystem, modify databases, send network requests, or change the host operating system.

Those capabilities come from tools supplied by the host.

That distinction is important for every agent integration, particularly on machines where a mistaken tool call could have real consequences.


Ollama agent compatibility

Ollama supports tool calling through its /api/chat endpoint and accepts tool schemas alongside messages. See the official Ollama tool-calling documentation:

https://docs.ollama.com/capabilities/tool-calling

ZynDwarf-1.1 was tested locally through this interface with the same host-controlled pattern.

Real smoke test

Tool supplied by the host:

{
  "type": "function",
  "function": {
    "name": "free_memory",
    "description": "Return the current available RAM in the system",
    "parameters": {
      "type": "object",
      "properties": {},
      "required": []
    }
  }
}

The model returned a structured call whose function name was free_memory for both published variants.

Measured local result

Variant Native tool-call smoke Observed
F16 1 / 1 message.tool_calls[0].function.name = free_memory
Q4_K_M 1 / 1 message.tool_calls[0].function.name = free_memory

Native Ollama tool-call smoke

This is a smoke test, not a claim that the model will correctly call every arbitrary tool on every prompt. The purpose is to verify that the released GGUF + Ollama packaging participates correctly in a real structured tool-call path.


Real local measurements

CPU throughput

The release was measured using llama-bench from llama.cpp on the development server.

Test conditions:

  • CPU: 2-vCPU Intel Xeon host;
  • threads: 2;
  • prompt tokens: 128;
  • generated tokens: 128;
  • llama.cpp build: b19cbe9, build 8;
  • measurement type: prompt processing (pp128) and token generation (tg128).

Raw result

Variant File size Prompt processing Generation
F16 676.25 MiB 221.00 ± 7.96 tok/s 16.51 ± 2.13 tok/s
Q4_K_M 216.41 MiB 281.80 ± 16.07 tok/s 38.52 ± 0.80 tok/s

Real ZynDwarf CPU throughput

Q4_K_M is approximately 68% smaller than F16 on disk and was materially faster in this CPU benchmark.

The throughput figures are machine-specific measurements. They are not a standardized model-quality benchmark and must not be compared to benchmark numbers quoted by another model vendor under a different GPU, CPU, context length, quantization, batch size, tokenizer, or runtime.

Q4 compression ratio

Using the measured GGUF sizes:

  • F16: 676.25 MiB
  • Q4_K_M: 216.41 MiB
  • reduction: about 68%

This is why Q4_K_M is the recommended default for smaller local machines, while F16 remains the reference artifact for fidelity comparisons.


Published comparison with other small models

The comparison below is intentionally split into two groups:

  1. our real local measurements, such as throughput and the Ollama tool-call smoke test;
  2. public benchmark numbers published by the upstream model authors.

These are not a single leaderboard.

A local 2-thread CPU throughput test and a benchmark score reported from an H100 cluster are measuring completely different things.

IFEval snapshot

IFEval is useful here because it focuses on instruction following rather than raw language-model memorization.

Liquid AI's current model card reports:

Model Approx. parameters IFEval
LFM2.5-350M 0.35B 76.96
Granite 4.0-H-350M 0.35B 61.27
Granite 4.0-350M 0.35B 53.48

Source:

https://huggingface.co/LiquidAI/LFM2.5-350M

The official SmolLM2 model card reports this compact-model comparison under its lighteval setup:

Model Approx. parameters IFEval
SmolLM2-360M-Instruct 0.36B 41.0
Qwen2.5-0.5B-Instruct 0.50B 31.6

Source:

https://huggingface.co/HuggingFaceTB/SmolLM2-360M-Instruct

Published IFEval comparison

Why ZynDwarf-1.1 is not assigned an IFEval number here

We did not run the IFEval benchmark itself in this release environment.

Instead, we ran real local runtime tests and smoke evaluations.

Assigning a made-up IFEval percentage by converting a different 1-to-1 smoke test into a leaderboard score would not be an evaluation. It would be decorative arithmetic.

The honest release therefore keeps the published upstream numbers separate from our own measurements.


Broader upstream comparison

Liquid AI's published model table is particularly useful because it places LFM2.5-350M directly beside models of a similar footprint.

Model GPQA Diamond MMLU-Pro IFEval BFCLv3 BFCLv4
LFM2.5-350M 30.64 20.01 76.96 44.11 21.86
LFM2-350M 27.58 19.29 64.96 22.95 12.29
Granite 4.0-H-350M 22.32 13.14 61.27 43.07 13.28
Granite 4.0-350M 25.91 12.84 53.48 39.58 13.73
Gemma 3 1B IT 23.89 14.04 63.49 16.61 7.17

Source:

https://huggingface.co/LiquidAI/LFM2.5-350M

These numbers should be interpreted as evidence about the upstream model family and its published evaluation, not as measured ZynDwarf-1.1 scores.

ZynDwarf-1.1 inherits the same underlying LFM2.5-350M architecture and tokenizer family, but it has been adapted with a separate behavioral training run.


Training details

Starting point

Base model:

LiquidAI/LFM2.5-350M

ZynDwarf-1.1 was not trained from scratch.

The release started from a repaired LFM2.5-350M-derived checkpoint and then applied a general agent-oriented LoRA adaptation.

Dataset

The intended training dataset contained:

  • 1,613 deduplicated sequences;
  • 1,503 sequences that successfully entered the encoder under the 768-token cap;
  • 308 examples containing real tool-oriented traces;
  • 62 multi-turn examples.

The behavioral mix covered:

  • normal conversation;
  • Spanish responses;
  • English responses;
  • programming;
  • debugging;
  • shell and Linux tasks;
  • structured output;
  • tool selection;
  • tool-call formatting;
  • multi-step work;
  • tool failure recovery;
  • cases where no tool should be used;
  • planning and task decomposition.

LoRA configuration

rank       = 8
alpha      = 16
dropout    = 0.05
targets    = q_proj, k_proj, v_proj
learning_rate = 1.5e-6
epochs     = 1
grad_accum = 4
max_length = 768

The adapter trained approximately 0.069% of the total model parameter count.

What this release does not claim

This release does not claim:

  • logit-level knowledge distillation;
  • 128K-token end-to-end training;
  • perfect tool use;
  • perfect reasoning;
  • broad factual reliability;
  • autonomous execution without a host;
  • parity with larger coding specialists.

The documentation is intentionally explicit about those boundaries.


Context length

There are three different concepts that are easy to accidentally collapse into one number.

1. Architectural maximum

The released Transformers configuration exposes a maximum position length of 128,000 tokens.

2. Training sequence cap

The fine-tuning run used a maximum encoded sequence length of 768 tokens.

3. Recommended runtime context

For practical local inference, the published examples use 32,768 tokens.

These values describe different parts of the system.

A model having 128K positional metadata does not mean that it was fine-tuned end-to-end on 128K-token conversations, nor does it mean that a 3.8 GiB RAM machine can comfortably process 128K tokens.

The release therefore treats 32K as the conservative runtime setting and 128K as an architectural configuration value.


Transformers

Installation

pip install -U transformers torch

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "itsZyn/ZynDwarf-1.1"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

messages = [
    {"role": "user", "content": "Hola, ¿qué puedes hacer?"}
]

inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt",
)

outputs = model.generate(
    inputs,
    max_new_tokens=128,
)

print(tokenizer.decode(outputs[0], skip_special_tokens=False))

Tool schemas

When the runtime and Transformers version support tool-aware chat templates, pass tool definitions through the tools= argument.

messages = [
    {"role": "user", "content": "Check the current memory."}
]

inputs = tokenizer.apply_chat_template(
    messages,
    tools=tools,
    add_generation_prompt=True,
    return_tensors="pt",
)

The host should parse and validate any generated tool call before execution.


llama.cpp

The GGUF files can be used directly with a current llama.cpp build that supports the LFM2 architecture.

F16

./llama-cli \
  -m ZynDwarf-1.1-f16.gguf \
  -c 32768 \
  -t 2 \
  --temp 0.35 \
  --top-k 40 \
  --top-p 0.9 \
  --repeat-penalty 1.05

Q4_K_M

./llama-cli \
  -m ZynDwarf-1.1-Q4_K_M.gguf \
  -c 32768 \
  -t 2 \
  --temp 0.35 \
  --top-k 40 \
  --top-p 0.9 \
  --repeat-penalty 1.05

For production agent work, prefer the runtime's structured chat API rather than manually concatenating raw control tokens unless you know exactly how that runtime handles the template.


Ollama

Pull the default Q4 release

ollama pull itsZyn/ZynDwarf-1.1:latest

Run

ollama run itsZyn/ZynDwarf-1.1:latest

F16 tag

ollama pull itsZyn/ZynDwarf-1.1:f16

Q4 tag

ollama pull itsZyn/ZynDwarf-1.1:q4_k_m

The default latest tag is the Q4_K_M build because that is the most practical variant for small local machines.

The f16 tag is the higher-fidelity reference build.

The q4_k_m tag is the compact build.


OpenAI-compatible local APIs

Ollama exposes an OpenAI-compatible endpoint. The model can therefore be used by applications that speak the OpenAI chat-completions protocol, subject to the specific client's tool-call support.

Example:

curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "itsZyn/ZynDwarf-1.1:latest",
    "messages": [
      {"role": "user", "content": "Hello"}
    ]
  }'

For structured agent use, the native Ollama /api/chat endpoint is preferable because it exposes the tools field directly:

https://docs.ollama.com/api/chat


Using ZynDwarf in an agent loop

A minimal conceptual Python loop looks like this:

from ollama import chat

messages = [
    {"role": "user", "content": "Check the current memory."}
]

def free_memory():
    # host-side implementation
    ...

tools = [free_memory]

while True:
    response = chat(
        model="itsZyn/ZynDwarf-1.1:latest",
        messages=messages,
        tools=tools,
    )

    messages.append(response.message)

    if not response.message.tool_calls:
        print(response.message.content)
        break

    for call in response.message.tool_calls:
        # validate the function name and arguments before execution
        if call.function.name == "free_memory":
            result = free_memory()
        else:
            raise ValueError("Unknown tool")

        messages.append({
            "role": "tool",
            "tool_name": call.function.name,
            "content": str(result),
        })

The example is intentionally host-centric. The model requests an operation; the application decides whether to execute it.


Security model

Treat every model-generated tool call as untrusted input.

At minimum, an agent runtime should validate:

  • tool name;
  • argument names;
  • argument types;
  • path restrictions;
  • allowed commands;
  • network destinations;
  • maximum execution time;
  • maximum output size;
  • permissions;
  • whether the requested action is actually necessary.

The model should not be the component that decides its own authorization policy.

For filesystem tools, prefer allowlisted project directories over unrestricted root access.

For shell tools, prefer narrowly defined functions over an unrestricted command interpreter.

For network tools, restrict destinations where possible.

For destructive operations, require an explicit policy or confirmation outside the model.

These practices are properties of the agent runtime rather than special powers granted by ZynDwarf itself.


Evaluation methodology

The release follows three separate evidence levels.

Level A: artifact validation

The GGUF files were loaded by llama.cpp and benchmarked successfully.

This verifies that the artifacts are structurally usable by the target runtime.

Level B: local runtime smoke tests

The F16 and Q4 Ollama packages were exercised through the actual /api/chat endpoint with a real tool schema.

Both returned a structured tool call for the test function.

This verifies the packaging path:

host tool schema -> Ollama -> ZynDwarf-1.1 -> structured tool call

Level C: upstream standardized benchmarks

Where standardized benchmark numbers are cited, they are labeled as upstream published results and linked to the originating model card.

No local smoke score is silently converted into IFEval, BFCL, MMLU, GPQA, GSM8K, HumanEval, or another standardized benchmark.


Limitations

ZynDwarf-1.1 is a small model.

At roughly 354.5M parameters, it has less representational capacity than much larger models. That affects:

  • difficult mathematics;
  • long-horizon planning;
  • broad factual recall;
  • deep debugging;
  • large repositories;
  • complicated multi-step reasoning;
  • unusual tool-use schemas;
  • long conversations with many state changes.

It can also make plausible mistakes.

Generated code still requires review.

Generated tool calls still require validation.

The model has no independent access to current information unless a host supplies a tool or retrieval system.

It is not intended for medical, legal, financial, security-critical, or other high-stakes decision making.


Known failure modes

A compact model can fail in ways that look surprisingly confident.

Common risk classes include:

Arithmetic drift

Short arithmetic problems can still produce incorrect intermediate reasoning, especially when the prompt requests an explanation instead of a single result.

Strict JSON failure

The model can occasionally add Markdown fences or explanatory text when the user requested JSON-only output. For applications requiring machine-readable output, validate the response and consider a structured-output-capable runtime.

Tool over- or under-use

The model may sometimes answer conceptually when a tool should have been used, or propose a tool when a direct answer would be sufficient. Agent runtimes should therefore implement a tool policy outside the model as well.

Context degradation

A large configured context window does not guarantee stable quality throughout the full window.

Quantization differences

Q4_K_M is faster and smaller, but it is not numerically identical to F16.

The F16 artifact should be treated as the reference model for fidelity investigations.


Why F16 and Q4 are both released

The two variants solve different deployment problems.

F16

Choose F16 when:

  • memory is available;
  • output fidelity matters more than disk size;
  • you are comparing future model changes;
  • you want a reference GGUF.

Q4_K_M

Choose Q4_K_M when:

  • RAM is constrained;
  • CPU inference matters;
  • local storage is limited;
  • you are building a lightweight agent on a modest device.

The current benchmark shows the expected trade-off clearly:

  • F16: 676.25 MiB, 16.51 tok/s generation;
  • Q4_K_M: 216.41 MiB, 38.52 tok/s generation;

on the exact test host and exact llama.cpp build documented above.


Reproducibility

To reproduce the release packaging, keep the following information together:

Base model: LiquidAI/LFM2.5-350M
LoRA rank: 8
LoRA alpha: 16
LoRA dropout: 0.05
Targets: q_proj, k_proj, v_proj
Learning rate: 1.5e-6
Epochs: 1
Gradient accumulation: 4
Max training sequence: 768
Encoded sequences: 1,503
Tool-oriented examples: 308
Multi-turn examples: 62

GGUF converter/runtime benchmark build: b19cbe9 (build 8)
F16 SHA256: ce452090b010d1e199fd7b39b2d65e5fcd8d7c2ceac52c38219f8b53e491c424
Q4 SHA256:  04a6ea0100b84a8687162db856f4c0b8042f54ad6b50a93b3c8132afc3499ae2

Always verify downloaded artifacts against SHA256SUMS before using them in a reproducible pipeline.


Version history

ZynDwarf-1.1

This release supersedes the public General Agent publication.

Highlights:

  • unified model name;
  • unified Hugging Face repository;
  • unified Ollama repository;
  • F16 + Q4_K_M GGUF distribution;
  • explicit agent metadata;
  • native tool-call documentation;
  • real Ollama structured tool-call smoke test;
  • refreshed llama.cpp throughput measurements;
  • published comparison against compact model peers;
  • reproducibility information;
  • explicit limitations and benchmark separation.

General Agent publication

The older General Agent naming was retired in favor of the versioned ZynDwarf-1.1 release line.


Credits and references

Base model

Liquid AI, LFM2.5-350M:

https://huggingface.co/LiquidAI/LFM2.5-350M

The upstream model card documents the base model architecture, supported languages, context configuration, tool-use format, and published evaluation results.

SmolLM2 comparison

Hugging Face, SmolLM2-360M-Instruct:

https://huggingface.co/HuggingFaceTB/SmolLM2-360M-Instruct

The model card provides published compact-model comparisons including SmolLM2-360M-Instruct and Qwen2.5-0.5B-Instruct.

Ollama tool calling

Official Ollama tool-calling documentation:

https://docs.ollama.com/capabilities/tool-calling

Ollama API

Official Ollama chat API:

https://docs.ollama.com/api/chat

Ollama model format

Official Ollama model import documentation:

https://docs.ollama.com/import

llama.cpp

https://github.com/ggml-org/llama.cpp


Citation

@misc{zyndwarf11,
  title  = {ZynDwarf-1.1},
  author = {Zyn Models and itsZyn},
  year   = {2026},
  note   = {Compact LFM2.5-350M based conversational and agent-oriented model}
}

Final release statement

ZynDwarf-1.1 is a compact local-first model for practical assistant and agent workflows.

It is intentionally small, intentionally measurable, and intentionally explicit about what has and has not been tested.

The F16 and Q4_K_M artifacts are released together.

The Ollama package is tested with real structured tool calls.

The throughput numbers are measured on the actual development CPU.

The external benchmark numbers are labeled as upstream numbers instead of being presented as if they were produced locally.

That is the release philosophy: small model, real runtime, real files, real measurements, and no imaginary leaderboard points.

Downloads last month
1,258
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for itsZyn/ZynDwarf-1.1

Quantized
(79)
this model
Quantizations
1 model