Kambo-v1

A 1.7B-parameter hybrid convolution/attention Mixture-of-Experts language model, trained from scratch. Two experts of sixteen are active per token, so a forward pass costs about 0.5B parameters.

This is a research model. You are free to use it for research, academic study, and evaluation. The only restriction is commercial use โ€” see the License.

Architecture

Total parameters 1,691,197,184
Active parameters per token 502,112,000
Layers 24
Hidden size 1024
Mixer 18 gated short-convolution layers + 6 grouped-query attention layers (at depths 3, 7, 11, 15, 19, 23)
Attention 16 query heads / 4 key-value heads, head dim 64
Feed-forward Mixture-of-Experts in every layer: 16 routed experts, top-2, plus one always-on shared expert
Expert hidden size 1152
Normalisation RMSNorm, pre-norm, with QK-norm on attention layers
Position encoding Rotary, theta 40,000
Context length 16,384
Vocabulary 151,936 (embeddings tied to the output head)
Precision bfloat16

Most layers are a double-gated short convolution rather than attention: the input is projected to three streams, two are multiplied together, passed through a causal depthwise convolution of kernel width 3, and gated by the third. This carries local context at a cost that does not grow with sequence length. Six attention layers, spaced evenly through the depth, carry the long-range dependencies. Because a convolution layer only needs to remember the last two columns of its input, the model's incremental-decoding state stays small as the context grows.

Every layer's feed-forward block is a Mixture-of-Experts. A router picks 2 of 16 experts per token; a shared expert runs for every token regardless. Routing is computed in float32 for stability and the published implementation is dropless โ€” no token is discarded at any capacity limit.

Training

Trained from scratch on 263.82B tokens in two phases:

Phase Tokens
Pretraining 214.84B
Post-training (supervised fine-tuning + reinforcement learning) 48.98B

Pretraining covers general text and a context-length extension to 16,384. Post-training covers supervised fine-tuning for chat, tool use and instruction following, followed by reinforcement learning with verifiable rewards on tool-calling and constraint-following tasks.

Benchmark contamination was controlled by construction: no BFCL, IFBench, or Multi-IF item appears in any training corpus at any stage.

Results

Model Params AA-Omniscience IFBench Multi-IF BFCLv3 BFCLv4
LFM2.5-8B-A1B 8B/A1B -24.70 56.47 79.93 64.79 49.73
Qwen3-30B-A3B-Thinking-2507 30.5B/3.3B -51.31 51.11 79.04 73.39 50.53
Gemma-4-26B-A4B-IT 26B/4B -62.07 47.25 82.06 68.87 55.87
gpt-oss-20b 21B/3.6B -49.17 58.65 76.64 62.52 49.88
Qwen3.5-4B 4B -51.53 50.38 67.43 71.06 54.01
Gemma-4-E4B-IT 8B -50.67 39.48 77.58 57.31 33.92
Gemma-4-E2B-IT 5.1B -72.00 33.53 69.70 56.44 31.91
Granite-4.0-H-Tiny 7B/A1B -75.50 21.28 59.00 56.89 28.52
Kambo-v1 1.7B/A0.5B -18.83 11.63 35.89 36.99* 37.27*

* BFCL figures for this model are AST accuracy averaged over the non-live and live splits; the peer figures are the overall leaderboard score.

On AA-Omniscience this model scores -18.83, the highest figure in this table โ€” but that is not a knowledge result. The index rewards declining over guessing, and the model answered only 6 of 600 questions. It is measuring silence, not recall. The honest reading of this table is that tool-call formatting is where this model comes closest to the field, and that everything else is well behind.

Usage

The model ships its own modeling code, so trust_remote_code=True is required on both the model and the tokenizer. It runs on CPU; in bfloat16 the weights are about 3.4 GB.

pip install torch transformers

Tested with transformers 5.14.1; the snippets below use argument spellings that also work on the 4.x series.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "VikramPal/kambo-v1"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,   # use torch.float32 on CPU if bfloat16 is slow
    device_map="auto",            # needs `accelerate`; omit to stay on CPU
)

messages = [{"role": "user", "content": "What is the capital of France?"}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=256, do_sample=True,
                         temperature=0.7, top_p=0.9)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:],
                       skip_special_tokens=True))

Prompt format

ChatML, and the packaged chat template reproduces the training format exactly:

<|im_start|>user
What is the capital of France?<|im_end|>
<|im_start|>assistant

Use apply_chat_template rather than building the string yourself. In particular, no system message is inserted when you do not supply one โ€” that matches how the model was trained, and prepending a default system prompt will push it off-distribution. Turns end with <|im_end|>.

Running on CPU

model = AutoModelForCausalLM.from_pretrained(
    model_id, trust_remote_code=True, torch_dtype=torch.float32
)

Batched generation requires left padding:

tokenizer.padding_side = "left"
batch = tokenizer(["Question one", "Question two"], return_tensors="pt", padding=True)
outputs = model.generate(**batch, max_new_tokens=64)

License

Released under the Kambo-v1 Research License โ€” non-commercial research and evaluation only. Commercial use of the model, its derivatives, or its outputs is prohibited. See LICENSE for the full terms, and NOTICE for third-party components distributed under their own licenses.

Citation

@misc{kambo_v1_2026,
  title  = {Kambo-v1: A 1.7B Hybrid Convolution-Attention Mixture-of-Experts Language Model},
  author = {Kamboj, Vikrampal},
  year   = {2026},
  note   = {Research model, non-commercial license},
  url    = {https://huggingface.co/VikramPal/kambo-v1}
}
Downloads last month
1,671
Safetensors
Model size
2B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support