Vertex 0.6 200M Base

A 198M-parameter decoder-only transformer (Qwen3 architecture), pretrained from scratch by Vertex Research on 20B tokens of filtered English web text (Ultra-FineWeb). It is the largest model in the Vertex 0.6 family and the strongest base we have trained so far. Apache 2.0, weights and tokenizer included.

This is a raw base model: plain-text completion only. No instruction tuning, no chat format, no safety tuning. Instruct and tool-calling variants built on this base are in progress.

Model details

Parameters 198.21M (tied embeddings)
Architecture Qwen3-based decoder-only transformer
Hidden size / layers 768 / 20
Attention 12 heads, 4 KV heads (GQA), head_dim 64
FFN 3072 (SwiGLU)
Context length 1024 (RoPE theta 10,000)
Vocabulary 32,768
Precision bf16 mixed-precision training, fp32 weights
Pretraining tokens 20B (about 100 tokens per parameter)

Benchmarks

Zero-shot, acc_norm, run with lm-eval-harness plus ArithMark-3. The "Avg" and "Intelligence Index" columns use the same weighting as every Vertex release (mean of HellaSwag, combined ARC, PIQA and ArithMark-3; the Index is chance-normalized with ArithMark weighted 0.65), so they are directly comparable to earlier Vertex numbers.

Every model below was run locally with the identical harness, tasks and formula on the same machine; no numbers are copied from other model cards. Pretrain token counts are as stated on each model's card ("?" means not disclosed). ARC is the mean of ARC-Easy and ARC-Challenge. BoolQ and SciQ are informational and not part of Avg / Index.

Model Params Pretrain tokens Avg Int. Index HellaSwag ARC (avg) PIQA ArithMark-3 BoolQ SciQ
SmolLM2-360M 362M 4T 60.14 41.87 56.4 53.0 72.0 59.2 61.9 86.5
LFM2-350M 354M 10T 52.38 32.79 49.0 52.9 69.6 38.0 64.4 89.4
SmolLM2-135M 135M 2T 48.40 26.76 43.0 44.0 68.5 38.1 60.1 78.3
SmolLM-135M 135M 600B 47.43 25.35 42.6 42.4 67.7 37.0 59.6 74.6
BananaMind-2-Pro 139M ? 47.21 24.92 42.8 40.7 67.6 37.7 58.7 76.2
Vertex 0.6 200M Base 198M 20B 46.02 23.30 40.5 40.8 67.0 35.8 56.8 70.7
Rose-1.5-Medium 99M ~100B 45.17 21.21 38.0 37.3 65.3 40.0 55.6 69.8
Rose-Pro 151M ? 44.71 20.78 38.1 37.8 65.1 37.8 60.5 68.8
Supra2-100M-Base 101M 30B 43.69 19.41 36.1 36.2 65.3 37.2 61.7 70.8
Mamba-130M 129M 300B 42.08 16.66 35.1 33.2 63.0 37.1 55.2 67.2
Vertex 0.6 100M Base 97M 10B 41.28 15.97 34.2 34.0 63.3 33.6 54.7 64.3
llama-160m 162M ? 40.23 14.71 34.6 29.6 64.0 32.7 61.7 60.2
GPT-2 (124M) 124M ~10B 39.84 13.36 31.3 31.5 61.5 35.1 50.1 64.1
OPT-125M 125M 180B 39.78 13.40 31.6 31.0 61.9 34.7 56.0 70.2
AMD-Llama-135m 134M 670B 38.02 10.45 27.1 28.8 60.5 35.7 61.8 46.7
TinyMistral-248M 248M ? 37.07 8.88 28.4 29.7 57.5 32.7 60.0 46.7
Pythia-160M 162M 300B 36.38 8.80 30.6 30.4 58.2 26.4 43.2 62.0

Reading the table: at 20B pretraining tokens this base lands above every model trained on 300B tokens or fewer, and above every model in the 100M to 160M range except the trillion-token SmolLM family and BananaMind-2-Pro. The gap to those is token exposure, not architecture: the models ahead of it saw 30x to 200x more data. Winogrande (53.0) and OpenBookQA (33.8) were also run; both sit near chance for every model in this size band and are omitted from the table.

The two 350M models are included as a ceiling reference only. BoolQ and ArithMark-3 are the columns with the most headroom for this base and are the targets of the continued-pretraining stage that follows.

Training

Single-run pretraining, about 11.6 days on one consumer GPU at ~20K tokens/s: cosine schedule, peak learning rate 6e-4 with 2,000 warmup steps, effective batch 64 sequences of 1024 tokens, AdamW, bf16 autocast. Training loss went from 9.1 to ~2.65.

Pretraining loss, 100M vs 200M

The curve is plotted against tokens seen, alongside the earlier 100M base (10B tokens); the 200M model sits below it at every point.

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM

repo = "VertexResearch/Vertex-0.6-200M-Base"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)

ids = tok("The capital of France is", return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=40, do_sample=True,
                     temperature=0.6, top_p=0.9, repetition_penalty=1.3)
print(tok.decode(out[0], skip_special_tokens=True))

Sampling (temperature around 0.6, top-p 0.9, repetition penalty around 1.3) is recommended; greedy decoding tends to loop at this model size.

Limitations

Fluent, well-structured English with reasonable common knowledge, but facts degrade quickly beyond the obvious: dates, names and numbers are often invented. It has seen no dedicated code data yet, so code completion is weak. No instruction following, no chat format, no safety tuning. Knowledge comes from web text and is not curated for accuracy. Not for production use; intended as a base for further training and for research on small models.

Downloads last month
286
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support