Instructions to use VikramPal/kambo-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VikramPal/kambo-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="VikramPal/kambo-v1", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("VikramPal/kambo-v1", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use VikramPal/kambo-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VikramPal/kambo-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VikramPal/kambo-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VikramPal/kambo-v1
- SGLang
How to use VikramPal/kambo-v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "VikramPal/kambo-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VikramPal/kambo-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "VikramPal/kambo-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VikramPal/kambo-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use VikramPal/kambo-v1 with Docker Model Runner:
docker model run hf.co/VikramPal/kambo-v1
Kambo-v1
A 1.7B-parameter hybrid convolution/attention Mixture-of-Experts language model, trained from scratch. Two experts of sixteen are active per token, so a forward pass costs about 0.5B parameters.
This is a research model. You are free to use it for research, academic study, and evaluation. The only restriction is commercial use โ see the License.
Architecture
| Total parameters | 1,691,197,184 |
| Active parameters per token | 502,112,000 |
| Layers | 24 |
| Hidden size | 1024 |
| Mixer | 18 gated short-convolution layers + 6 grouped-query attention layers (at depths 3, 7, 11, 15, 19, 23) |
| Attention | 16 query heads / 4 key-value heads, head dim 64 |
| Feed-forward | Mixture-of-Experts in every layer: 16 routed experts, top-2, plus one always-on shared expert |
| Expert hidden size | 1152 |
| Normalisation | RMSNorm, pre-norm, with QK-norm on attention layers |
| Position encoding | Rotary, theta 40,000 |
| Context length | 16,384 |
| Vocabulary | 151,936 (embeddings tied to the output head) |
| Precision | bfloat16 |
Most layers are a double-gated short convolution rather than attention: the input is projected to three streams, two are multiplied together, passed through a causal depthwise convolution of kernel width 3, and gated by the third. This carries local context at a cost that does not grow with sequence length. Six attention layers, spaced evenly through the depth, carry the long-range dependencies. Because a convolution layer only needs to remember the last two columns of its input, the model's incremental-decoding state stays small as the context grows.
Every layer's feed-forward block is a Mixture-of-Experts. A router picks 2 of 16 experts per token; a shared expert runs for every token regardless. Routing is computed in float32 for stability and the published implementation is dropless โ no token is discarded at any capacity limit.
Training
Trained from scratch on 263.82B tokens in two phases:
| Phase | Tokens |
|---|---|
| Pretraining | 214.84B |
| Post-training (supervised fine-tuning + reinforcement learning) | 48.98B |
Pretraining covers general text and a context-length extension to 16,384. Post-training covers supervised fine-tuning for chat, tool use and instruction following, followed by reinforcement learning with verifiable rewards on tool-calling and constraint-following tasks.
Benchmark contamination was controlled by construction: no BFCL, IFBench, or Multi-IF item appears in any training corpus at any stage.
Results
| Model | Params | AA-Omniscience | IFBench | Multi-IF | BFCLv3 | BFCLv4 |
|---|---|---|---|---|---|---|
| LFM2.5-8B-A1B | 8B/A1B | -24.70 | 56.47 | 79.93 | 64.79 | 49.73 |
| Qwen3-30B-A3B-Thinking-2507 | 30.5B/3.3B | -51.31 | 51.11 | 79.04 | 73.39 | 50.53 |
| Gemma-4-26B-A4B-IT | 26B/4B | -62.07 | 47.25 | 82.06 | 68.87 | 55.87 |
| gpt-oss-20b | 21B/3.6B | -49.17 | 58.65 | 76.64 | 62.52 | 49.88 |
| Qwen3.5-4B | 4B | -51.53 | 50.38 | 67.43 | 71.06 | 54.01 |
| Gemma-4-E4B-IT | 8B | -50.67 | 39.48 | 77.58 | 57.31 | 33.92 |
| Gemma-4-E2B-IT | 5.1B | -72.00 | 33.53 | 69.70 | 56.44 | 31.91 |
| Granite-4.0-H-Tiny | 7B/A1B | -75.50 | 21.28 | 59.00 | 56.89 | 28.52 |
| Kambo-v1 | 1.7B/A0.5B | -18.83 | 11.63 | 35.89 | 36.99* | 37.27* |
* BFCL figures for this model are AST accuracy averaged over the non-live and live splits; the peer figures are the overall leaderboard score.
On AA-Omniscience this model scores -18.83, the highest figure in this table โ but that is not a knowledge result. The index rewards declining over guessing, and the model answered only 6 of 600 questions. It is measuring silence, not recall. The honest reading of this table is that tool-call formatting is where this model comes closest to the field, and that everything else is well behind.
Usage
The model ships its own modeling code, so trust_remote_code=True is required
on both the model and the tokenizer. It runs on CPU; in bfloat16 the weights are
about 3.4 GB.
pip install torch transformers
Tested with transformers 5.14.1; the snippets below use argument spellings
that also work on the 4.x series.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "VikramPal/kambo-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16, # use torch.float32 on CPU if bfloat16 is slow
device_map="auto", # needs `accelerate`; omit to stay on CPU
)
messages = [{"role": "user", "content": "What is the capital of France?"}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, do_sample=True,
temperature=0.7, top_p=0.9)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:],
skip_special_tokens=True))
Prompt format
ChatML, and the packaged chat template reproduces the training format exactly:
<|im_start|>user
What is the capital of France?<|im_end|>
<|im_start|>assistant
Use apply_chat_template rather than building the string yourself. In
particular, no system message is inserted when you do not supply one โ that
matches how the model was trained, and prepending a default system prompt will
push it off-distribution. Turns end with <|im_end|>.
Running on CPU
model = AutoModelForCausalLM.from_pretrained(
model_id, trust_remote_code=True, torch_dtype=torch.float32
)
Batched generation requires left padding:
tokenizer.padding_side = "left"
batch = tokenizer(["Question one", "Question two"], return_tensors="pt", padding=True)
outputs = model.generate(**batch, max_new_tokens=64)
License
Released under the Kambo-v1 Research License โ non-commercial research and evaluation only. Commercial use of the model, its derivatives, or its outputs is prohibited. See LICENSE for the full terms, and NOTICE for third-party components distributed under their own licenses.
Citation
@misc{kambo_v1_2026,
title = {Kambo-v1: A 1.7B Hybrid Convolution-Attention Mixture-of-Experts Language Model},
author = {Kamboj, Vikrampal},
year = {2026},
note = {Research model, non-commercial license},
url = {https://huggingface.co/VikramPal/kambo-v1}
}
- Downloads last month
- 1,671