How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf CodeStrux-Tech/tac-1-gguf:Q5_K_M
# Run inference directly in the terminal:
llama cli -hf CodeStrux-Tech/tac-1-gguf:Q5_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf CodeStrux-Tech/tac-1-gguf:Q5_K_M
# Run inference directly in the terminal:
llama cli -hf CodeStrux-Tech/tac-1-gguf:Q5_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf CodeStrux-Tech/tac-1-gguf:Q5_K_M
# Run inference directly in the terminal:
./llama-cli -hf CodeStrux-Tech/tac-1-gguf:Q5_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf CodeStrux-Tech/tac-1-gguf:Q5_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf CodeStrux-Tech/tac-1-gguf:Q5_K_M
Use Docker
docker model run hf.co/CodeStrux-Tech/tac-1-gguf:Q5_K_M
Quick Links

tac-1-gguf — GGUF format for llama.cpp and Ollama

Overview

tac-1-gguf provides GGUF files quantized from CodeStrux-Tech/tac-1 using llama.cpp commit 67776ea.

file size sha256
tac-1-Q5_K_M.gguf 2.7 GB 1c919ee74addda61bb6490195934b7c06b1a1793051e6119a2859d045e5b0634
tac-1-f16.gguf 7.5 GB 5ba9ce916778026da6a4970b41fe4f230b9a7c30c0f0437d3630724c7c94feea

Serving

Ollama

ollama run hf.co/CodeStrux-Tech/tac-1-gguf:Q5_K_M

Ollama registers an HF-pulled model under the full hf.co/...:Q5_K_M name. To use it with the tico client (default model name tac-1), alias it once with ollama cp hf.co/CodeStrux-Tech/tac-1-gguf:Q5_K_M tac-1, or set TICO_OLLAMA_MODEL="hf.co/CodeStrux-Tech/tac-1-gguf:Q5_K_M".

llama.cpp server

llama-server -m tac-1-Q5_K_M.gguf --jinja

The server is OpenAI-compatible at /v1. Client env: OLLAMA_BASE_URL=http://localhost:8080/v1 TICO_OLLAMA_MODEL=tac-1.

Chat template warning

The chat template is ChatML with EOS id 151645 (<|im_end|>). There are 0 think references in the template. The template is byte-identical across the merged bf16, FP8, and GGUF builds. When using llama.cpp or Ollama, ensure the server applies the ChatML template correctly — the --jinja flag on llama-server enables this.

Training data attribution

Contains information from OpenStreetMap (https://www.openstreetmap.org/copyright), which is made available under the Open Database License (ODbL) 1.0. © OpenStreetMap contributors.

For full training details, architecture, evaluation, and limitations, see CodeStrux-Tech/tac-1.

tac-1 is a derivative work of Qwen/Qwen3-4B-Instruct-2507, Copyright 2024 Alibaba Cloud, licensed under the Apache License, Version 2.0. The upstream LICENSE is included in this repository.

Downloads last month
168
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for CodeStrux-Tech/tac-1-gguf

Quantized
(3)
this model