Instructions to use redhamohamed/naim with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use redhamohamed/naim with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="redhamohamed/naim", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("redhamohamed/naim", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use redhamohamed/naim with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf redhamohamed/naim:Q4_K_M # Run inference directly in the terminal: llama cli -hf redhamohamed/naim:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf redhamohamed/naim:Q4_K_M # Run inference directly in the terminal: llama cli -hf redhamohamed/naim:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf redhamohamed/naim:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf redhamohamed/naim:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf redhamohamed/naim:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf redhamohamed/naim:Q4_K_M
Use Docker
docker model run hf.co/redhamohamed/naim:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use redhamohamed/naim with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "redhamohamed/naim" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redhamohamed/naim", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/redhamohamed/naim:Q4_K_M
- SGLang
How to use redhamohamed/naim with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "redhamohamed/naim" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redhamohamed/naim", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "redhamohamed/naim" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redhamohamed/naim", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use redhamohamed/naim with Ollama:
ollama run hf.co/redhamohamed/naim:Q4_K_M
- Unsloth Desktop
- Pi
How to use redhamohamed/naim with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf redhamohamed/naim:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "redhamohamed/naim:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use redhamohamed/naim with Docker Model Runner:
docker model run hf.co/redhamohamed/naim:Q4_K_M
- Lemonade
How to use redhamohamed/naim with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull redhamohamed/naim:Q4_K_M
Run and chat with the model
lemonade run user.naim-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use redhamohamed/naim with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf redhamohamed/naim:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default redhamohamed/naim:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use redhamohamed/naim with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf redhamohamed/naim:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "redhamohamed/naim:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Naim 9B · v2
Naim is an autonomous AI coding agent created by ABDESSEMED Mohamed (redhamohamed).
Lineage: Mimo → Naim
Mimo (redhamohamed/mimo-v5-gguf · 8B · by ABDESSEMED Mohamed)
│ identity, name, spirit: a private, local assistant that belongs to its creator
▼
Naim (redhamohamed/naim · Naim architecture · 9B · autonomous coding agent)
Naim is the successor of Mimo, Mohamed's earlier model (not related to Xiaomi's MiMo). Naim carries Mimo's identity and purpose forward as an autonomous coding agent.
Naim architecture
| Parameters | ~9B |
| Context | up to 262,144 tokens |
| Modes | reasoning (<think>) and direct answer, tool calling, vision (with naim-mmproj-f16.gguf) |
| Languages | French, English and many others |
model_type |
naim (NaimForConditionalGeneration) |
How Naim works
Naim follows an agent loop on every task:
- Understand: restate the goal, read the relevant code, list constraints.
- Plan: split the work into small verifiable steps.
- Implement: clean, idiomatic code that matches the existing style.
- Verify: run tests or scripts, read errors, fix.
- Report: what changed, what was verified, what's left.
Files
| File | Use |
|---|---|
*.safetensors + *_naim.py |
Transformers (bf16, trust_remote_code=True) |
naim-Q4_K_M.gguf |
LM Studio / Ollama / llama.cpp: recommended (~5.6 GB) |
naim-Q8_0.gguf |
Higher quality (~9.5 GB) |
naim-mmproj-f16.gguf |
Vision projector for LM Studio / llama.cpp (lets Naim see images) |
Usage
Ollama
ollama run hf.co/redhamohamed/naim:Q4_K_M
LM Studio: search redhamohamed/naim and download naim-Q4_K_M.gguf (+ naim-mmproj-f16.gguf for images).
llama.cpp (with vision):
llama-server -m naim-Q4_K_M.gguf --mmproj naim-mmproj-f16.gguf
Transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("redhamohamed/naim", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("redhamohamed/naim", trust_remote_code=True, dtype=torch.bfloat16)
msgs = [{"role": "user", "content": "Qui es-tu ?"}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt", return_dict=True)
out = model.generate(**ids, max_new_tokens=512)
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))
Pass enable_thinking=False to apply_chat_template for direct answers without the reasoning phase.
Recommended sampling for code: temperature 0.3, top_p 0.95, top_k 20, presence_penalty 0 (a presence penalty breaks code: braces and names must repeat).
What's new in v2
- No more canned self-introductions: Naim answers "merci" like a normal assistant and only presents itself when asked.
- Answers distilled to keep the full coding skill (complete, compilable code with every import).
- Knows its identity (creator, Mimo lineage, Naim architecture) without inventing details.
License
Apache 2.0. Naim is a fine-tune of an open-weight Apache-2.0 base model; see the LICENSE file.
- Downloads last month
- -