Mark 1x-9B

Answers in interfaces, not paragraphs.

Ask most models for an explanation and you get prose. Mark 1x-9B decides when a question deserves a 3D scene you can rotate, a chart with the real numbers in it, or a quiz that marks itself — then writes one that parses and renders on the first try, with no schema and no constrained decoding.

Before training Mark 1x-9B
Directive valid on the first try 0% 84%
Restraint (answers in prose when no interface helps) 75% 87%
Multi-turn edits to an artifact on screen 5/14 12.3/14

Everything else — code, mathematics, tool calling, the 262,144-token context, 201 languages, thinking mode — is inherited and deliberately left alone. AIME 2026 83.3 · AIME 2025 78.3 · MMLU-Pro 80.6 · GPQA Diamond 79.4.

Built on Qwen/Qwen3.5-9B (Apache-2.0) · Method LoRA r=32, merged to bf16 · Context 262,144 tokens · Licence OpenMDW-1.0 · Version v1.1b

The directive contract (SPEC.md) · Launch post · Try the assistant · @SaanoraLabs on X

Read Benchmarks for the full protocol and Limitations before you depend on this model. Every number here is self-measured (2026-09-03) with the harness and raw generations published alongside the weights — check them rather than trust them.

What this is

Mark 1x-9B is a nine-billion-parameter model trained to emit and reason about Saanora's directive DSL — a small set of <kind>{...JSON...}</kind> tags (scene, chart, quiz, flashcards, simulation, explore, steps, support, plus the file/text kinds pdf, docx, xlsx, pptx, copy — 13 kinds in total) that a frontend renderer turns into interactive UI: 3D scenes, charts, quizzes, flashcards, physics simulations, zoomable concept trees, step walkthroughs, and generated documents.

The one thing this model does: decide when to emit a directive and produce one that parses and renders on the first try. Measured on schema-free (unconstrained) decoding:

Before training Mark 1x-9B v1.1b
Directive validity (schema-free) 0% 84%
Restraint (correctly emits no directive when none helps) 75% 87%
Multi-turn artifact edits (out of 14) 5/14 12.3/14
Native tool calling preserved preserved

Without this model — or without knowing the DSL exists — a caller who sends a directive-shaped request to an untrained model gets raw, usually-malformed <scene>{...}</scene> text in the chat output and reasonably concludes the model is broken. It isn't broken; it's speaking a contract you need to read once. Jump to "How to use the directives."

Everything other than the DSL — code, math, tool-calling, the 262K context, 201 languages, thinking mode — is inherited, not retrained, and the job of the training was to leave those alone. It mostly did (see Benchmarks), but not perfectly (see Limitations).

Quickstart

transformers

The release checkpoint is the text-only merge: Qwen3_5ForCausalLM (model_type: qwen3_5_text), 32 layers, no vision encoder — see Vision below. Load it with transformers as a causal LM; image content parts are not accepted.

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_id = "Saanora/mark-1x-9b"

tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

messages = [
    {"role": "system", "content": "You are Mark, Saanora's assistant."},
    # No mention of a chart, a quiz or a scene. Deciding which one
    # this question deserves is the model's job, and the whole point.
    {"role": "user", "content": "Quiz me on the water cycle."},
]

# enable_thinking must be passed EXPLICITLY. Leaving it to the default can
# silently produce an empty <think></think> block that looks fine but is a
# known failure mode for this template family (see Limitations).
text = tok.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=True
)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=4096)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

vLLM (recommended for serving)

Exact flags, reproduced from this project's api/serve/FLAGS.md (every flag there is cited to either the live vLLM docs or a benchmark run that actually used it):

VLLM_USE_FLASHINFER_SAMPLER=0 \
HF_HUB_ENABLE_HF_TRANSFER=1 \
TOKENIZERS_PARALLELISM=false \
PYTHONUTF8=1 \
vllm serve Saanora/mark-1x-9b \
  --served-model-name mark-1x-9b \
  --max-model-len 262144 \
  --dtype bfloat16 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --generation-config vllm \
  --structured-outputs-config.backend auto \
  --structured-outputs-config.enable_in_reasoning=False \
  --enable-prefix-caching \
  --max-num-seqs 64 \
  --max-num-batched-tokens 16384 \
  --gpu-memory-utilization 0.90

Notes:

  • --language-model-only is harmless here: the release checkpoint is text-only, so there is no vision tower to strip. This project's own serving config passes it for parity with the benchmark runs.
  • --structured-outputs-config.enable_in_reasoning must be False. With True, every request carrying response_format returns content: null with the answer buried in message.reasoning and reasoning_tokens misreported as 0 (measured 2026-09-07, vLLM 0.28.0). Any OpenAI-compatible client reading content gets nothing.
  • --max-num-seqs 64 at the full --max-model-len 262144 was load-tested on one A100-80GB (2026-09-07): 48 concurrent requests, no OOM, 2,375 output tok/s aggregate. It is a hybrid (8 full-attention + 24 linear-attention layers), so the growing KV cache is ~32 KiB/token — 8 GiB per full-length sequence. On smaller cards reduce --max-model-len or --max-num-seqs; a 24 GB card cannot hold the weights plus one full-length sequence.
  • Always pass thinking mode explicitly per request (extra_body={"chat_template_kwargs": {"enable_thinking": true}}), same reasoning as above.
  • Reasoning is returned in choices[0].message.reasoning (older vLLM: .reasoning_content), separate from .content.
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

resp = client.chat.completions.create(
    model="mark-1x-9b",
    messages=[
        {"role": "system", "content": "You are Mark, Saanora's assistant."},
        {"role": "user", "content": "Explain how binary search works."},
    ],
    extra_body={"chat_template_kwargs": {"enable_thinking": True}},
    max_tokens=4096,
)
msg = resp.choices[0].message
print(msg.reasoning)   # the thinking, in its own field
print(msg.content)     # prose, then the directive last

SGLang

Supported upstream — SGLang's model registry carries qwen3_5.py and qwen3_5_text.py, so this checkpoint needs nothing special. Verify parser names against your SGLang version.

python -m sglang.launch_server \
  --model-path Saanora/mark-1x-9b \
  --served-model-name mark-1x-9b \
  --context-length 262144 \
  --dtype bfloat16 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen25

Engine support at a glance

Because this is a merged LoRA over an unmodified architecture, every engine that runs the this architecture runs this model too, with no porting and no upstream patch.

Engine Status
vLLM Verified here — every benchmark below was produced by vLLM 0.28.0 serving this checkpoint
SGLang Supported upstream (in the model registry)
transformers Verified here
llama.cpp / Ollama / LM Studio, TensorRT-LLM, TGI Not tested by us — no claim made

How to use the directives

The full contract is SPEC.md, published alongside this checkpoint. The short version:

1. Envelope: <kind>{ …strict JSON… }</kind>, where kind is one of chart, scene, quiz, flashcards, simulation, explore, steps, support (JSON body) or pdf, docx, xlsx, pptx, copy (document/text kinds — the first three of those five are free text, not JSON). Prose before and after the block is normal markdown.

2. Ordering is load-bearing: reasoning → prose → directive, always, directive last. A well-formed turn looks like:

<think>
The user wants to see benzene's structure. A 3D scene with a ring of six
carbons and alternating double bonds communicates this better than prose.
</think>

Benzene is a six-carbon aromatic ring. The delocalised electrons sit above
and below the plane, which is why it's drawn with a circle inside the hexagon.

<scene>{"title":"Benzene","environment":"studio","objects":[
  {"type":"ring","radius":2,"tube":0.15,"color":"#888","rotation":[1.5708,0,0]},
  {"type":"sphere","position":[2,0,0],"radius":0.3,"color":"#333","label":"C"},
  {"type":"sphere","position":[1,1.73,0],"radius":0.3,"color":"#333","label":"C"}
]}</scene>

The reply to "Explain how binary search works" has the same shape — prose, then one directive the renderer can draw. Illustrative, not a captured transcript; the exact field names are in SPEC.md:

Binary search halves the search range at every step, so a list of a
million items is exhausted in about twenty comparisons.

<steps>{"title":"Binary search for 23","steps":[
  {"label":"lo=0 hi=6 mid=3","detail":"12 < 23, search the right half"},
  {"label":"lo=4 hi=6 mid=5","detail":"23 == 23, found at index 5"}
]}</steps>

3. Restraint matters as much as emission. A directive is for something spatial, quantitative, procedural, or drillable. "What's the capital of France?" should get plain prose, not a directive. Over-emitting is a failure mode we measured (see the restraint numbers above) — expect occasional both-direction misses, not perfection.

4. Failure is always graceful on the render side, but the JSON must still be valid to be worth anything: every block is parsed with try { JSON.parse } catch { return {error} }. A malformed body renders a "couldn't render" card, not a crash — but it's also not useful output. Unknown/extra JSON keys are silently ignored by the renderer, so a slightly-too-generous completion is harmless; a syntactically broken one isn't.

Guaranteeing valid directives: guided decoding

Unconstrained, this model emits schema-free-valid directives 84% of the time (see above) — usable, not perfect. Two ways to do better, both shipped in api/grammars/:

  • response_format + a per-kind JSON Schema (recommended when your app already knows the kind). Point response_format at grammars/schemas/strict/<kind>.schema.json (e.g. scene.schema.json, chart.schema.json) — strict/ uses additionalProperties: false throughout, which is what a grammar compiler needs (an open key set has no finite grammar). This constrains only the model's final, post-thinking output, so what you get back is the bare JSON body, not the <kind>...</kind> tags or the surrounding prose — you supply the envelope yourself. Use this when your product already decided "give me a scene for X" (e.g. a "generate a chart" button) rather than letting the model decide whether/what to emit.

    import json
    scene_schema = json.load(open("grammars/schemas/strict/scene.schema.json"))
    
    resp = client.chat.completions.create(
        model="mark-1x-9b",
        messages=[{"role": "user", "content": "A 3D scene of benzene, 6-8 objects."}],
        extra_body={"chat_template_kwargs": {"enable_thinking": True}},
        response_format={"type": "json_schema", "json_schema": {"name": "scene", "schema": scene_schema}},
    )
    
  • Full-envelope grammar (grammars/envelope.ebnf), for the harder guarantee "valid envelope syntax, and if a directive appears its JSON at least parses" with the kind left open to the model — this is what lets the model keep deciding whether prose alone is the right answer. This is the only option that can enforce the think → prose → directive ordering itself, since response_format never sees the tag or the surrounding prose. grammars/chart.ebnf is a worked, fully-typed proof for one kind. Caveat inherited from that file: the exact EBNF operator subset vLLM's grammar backend accepts was not independently re-verified against a live server (only a minimal example is confirmed in vLLM's own docs) — smoke-test before depending on this path in production. Pass it as extra_body={"structured_outputs": {"grammar": open("grammars/envelope.ebnf").read()}} (nested under structured_outputs, not flat — confirmed against a live docs fetch during this project).

  • Validating output that already exists (a QA pass, a logging pipeline, a flagged-completion review) — use grammars/schemas/permissive/<kind>.schema.json instead (additionalProperties: true, same required-field/type rules). It matches what the renderer actually tolerates; strict/ will reject a handful of real, renderer-accepted completions over a harmless extra field, so don't use strict/ as a post-hoc validity check.

Note the strict/scene.schema.json validator is deliberately tighter than the renderer: the renderer's only check is Array.isArray(objects), so {"objects": []} or a scene of one grey unlabeled cube "validates" there. The strict schema additionally requires objects.length >= 1 (capped at 400), and per-type required fields (sphere needs radius, box needs size, link/arrow need from/to, etc.) — use it, not just "does it parse," as your bar for "good."

Benchmarks

Protocol ("card protocol," directly comparable to published cards in this class): thinking mode ON; the published sampling settings for this architecture ("general": T=1.0, top-p=0.95, top-k=20, presence=1.5; "coding": T=0.6, presence=0); 32,768-token output cap (32k is the published recommendation for most queries, 81,920 for the hardest); single sample per prompt except math (avg@2) and GPQA (avg@4). Self-measured — as every model card is — but raw outputs, per-prompt scores, and harness code are published so anyone can re-score.

Benchmark Mark 1x-9B
IFEval (prompt-strict) 84.7
IFBench (prompt-loose) 56.0
MMLU-Pro (700-Q stratified subset) 80.6
GPQA Diamond (avg@4) 79.4
AIME 2025 (avg@2) 78.3
AIME 2026 (avg@2) 83.3
HMMT Feb 2026 (avg@2) 63.6

LiveCodeBench v6 is deliberately not reported as a headline number. Under the same 32k-token cap, 74 of 175 problems were truncated with no code emitted at all and scored zero — the raw number (47.4) would misrepresent the model's actual coding ability rather than measure it; it's a budget artifact, not a capability finding. A partial 81,920-token rerun (which the card's own guidance recommends for hard problems) reached 123/175 before being stopped and is archived but incomplete. Full numbers and the reasoning are in RESULTS.md for anyone who wants them; they are intentionally withheld from this card.

Peer context (dense, ≤14B, published card numbers)

Benchmark Mark 1x-9B Gemma 4 12B Ornith 1.5-9B Nemotron Nano 9B v2 ZAYA1-8B
9B 12B 9B 9B v2 8B MoE
Instruction following
IFEval — prompt-level strict 84.7 97.2 — 90.3‡ —
IFBench — prompt-level loose 56.0 — — — —
Knowledge and reasoning
MMLU-Pro* 80.6 77.2 — — —
GPQA Diamond — avg@4 79.4 78.8 86.4 64.0 —
Mathematics
AIME 2025 — avg@2, no tools 78.3 — — 72.1 —
AIME 2026 — avg@2, no tools 83.3 77.5 — — 89.1
HMMT Feb 2026 — avg@2, no tools 63.6 — — — 71.6
Directive output
Valid on the first try — schema-free 84.0 — — — —
Restraint — prose when no interface helps 87.0 — — — —
Multi-turn artifact edits — out of 14 12.3 — — — —

* 700-question stratified subset. ‡ instruction-level strict, not prompt-level.

An em-dash means that model publishes no number for the benchmark, which is not the same as scoring zero. Our rows are self-measured with every raw generation published; peer rows are vendor-reported under their own protocols, so read the comparison as indicative rather than like-for-like. LiveCodeBench is absent by decision — see the Benchmarks section above for why.

Reading it honestly: on the rows we ran, Mark 1x-9B sits with the Gemma 4 12B / Nemotron Nano 9B group on knowledge and math (GPQA 79.4 above Nemotron Nano 9B v2, Gemma 4 12B, and every 30B-class MoE we compared against except two; AIME 2025 level with Ministral 3 8B). ZAYA1-8B leads both maths rows in the table above — AIME 2026 and HMMT Feb 2026 — on a fraction of the active parameters, and Ornith 1.5-9B leads GPQA Diamond at the same size; both are stated here because both are visible in the same table. Coding (LiveCodeBench, SWE-bench, Terminal-Bench) is where agentic-coding-focused labs (Ornith, Poolside, GLM) invested and where this model is weakest — that's harness work ahead, not a claim made here. Full peer tables (20B–35B MoE tier, wide single-row view, sources) are in RESULTS.md.

Put a harness in front of the tools

Separate row, not comparable to the card protocol above: thinking ON, agentic loop, up to 10 tool rounds, python / web search / page fetch / a public-test runner, no API keys.

Benchmark Tool-free With tools What happened
IFEval (prompt-strict) 84.7 77.3 formatting constraints broken by the loop
MMLU-Pro subset 80.6 71.9 noisy raw web snippets misled the model (167/700 searched)
AIME 2026 (avg@2) 83.3 55.0 24/60 runs ended with no boxed answer
HMMT Feb 2026 (avg@2) 63.6 42.4 34/66 runs ended with no boxed answer

A 9B model without agentic post-training mostly ignores tools, and when it does use them, pays in tokens, formatting, and (with raw web snippets) accuracy. Tools belong behind a harness that decides when to invoke them — execute-and-repair for code, a final formatting pass, curated retrieval — not exposed as free-rein tool_choice:"auto" to an end user. See Limitations.

Where the adapter actually landed

This is a hybrid architecture: 32 transformer layers with full_attention_interval: 4, meaning 8 layers use full attention and the other 24 use linear (Mamba-style) attention. That shape decides where a LoRA can reach, and the result is worth publishing on its own:

  • The LoRA adapter covers mlp.gate_proj / mlp.up_proj / mlp.down_proj in all 32 layers.
  • It covers self_attn.q_proj / k_proj / v_proj / o_proj in only the 8 full-attention layers — the 24 linear-attention blocks use different projection names entirely, so the adapter simply doesn't touch attention-equivalent weights there.

In plain terms: this adapter does not cover "all projections." Three-quarters of the network's attention-equivalent path was never adapted, so the directive behaviour runs almost entirely through the MLP path (present in every layer) plus attention in the 8 full-attention layers. This is very likely part of why the regressions below concentrate where they do rather than being spread evenly, though this card does not claim a proven causal link.

Scope: text in, text out

This checkpoint is text-only: Qwen3_5ForCausalLM, 32 layers, no vision or audio encoder, and no image or audio generation. Image input is rejected outright rather than silently degraded, so vLLM returns HTTP 400 (mark-1x-9b is not a multimodal model, measured 2026-09-07). A request to draw something becomes a <scene> or <chart> directive your renderer draws, which is the point of the model.

What we have not measured

So nothing here is mistaken for a claim, this is the state of the evidence. Not measured by this project: image and video input (no MMMU-Pro or any vision benchmark was run, on either checkpoint); long-context behaviour anywhere near the 262,144-token limit; LiveCodeBench under a sufficient output budget (our run capped generation at 32k tokens, which zeroed 74 of 175 problems before any code was emitted, so no coding number is published here); SWE-bench Verified and Terminal-Bench; multilingual quality beyond spot checks; and every OpenAI-API surface claim against our own server rather than a hosted copy. A missing number means the evaluation was not run.

Known limits, and how to work around them

  • Exact-text constraints are the weak spot, and the shortfall is concentrated rather than spread: repeat 0%, custom 27%, words 42%, ratio 45%, versus format / count / sentence constraints at 71–78%. The pattern says targeted replay data rather than a lost capability, and it is first on the roadmap. As shipped, don't depend on it for exact repeat-N-times, exact-word-count or exact-ratio instructions.
  • Give it a tool harness, not free rein. Unsupervised tool use costs accuracy on every benchmark measured (see table above: IFEval 84.7→77.3, MMLU-Pro 80.6→71.9 with raw web snippets, both math benchmarks down 20+ points). If you expose tools to this model, put a harness in front that decides when to call them — don't hand it tool_choice:"auto" and walk away.
  • Always set a system prompt. Identity is barely trained, by design. Only 2–3 of 1,396 training rows even touch an identity-shaped question. With no system prompt it answers "what are you" from its pretraining priors rather than a trained-in persona. Name the assistant in your system prompt (see the Quickstart examples above) if you need consistent identity behavior — this is a deployment fix, not something the weights alone guarantee.
  • One training row taught a non-renderable scene object type. A small amount of label noise made it into the training set; if you see a <scene> silently render fewer objects than it describes, this is a known, minor contributor.
  • GPQA Diamond at 79.4 (avg@4) is a real, small cost of teaching the directive skill, not a protocol artifact. The spread across 4 passes was only 1.5 points, and 136/198 questions were answered identically in all four passes.
  • LiveCodeBench is deliberately withheld from this card (see Benchmarks) — treat this model as unevaluated for competitive coding, not as scoring 47.4.
  • Vision is not in this release (see above) — image input returns HTTP 400. Don't build a product feature on it.
  • This is a 9B model. It wins on the one thing it was built for and holds its own elsewhere in its class. Compare it with the small, cheap tier it belongs to, not with the frontier.

License and attribution

  • Weights are released under OpenMDW-1.0 (Open Model, Data and Weights Licence 1.0). The full license text ships as LICENSE alongside the weights.
  • This model is a derivative work of Qwen/Qwen3.5-9B, licensed under the Apache License 2.0. Required attribution, copyright notice, and statement of modification are in NOTICE — read it alongside this card; it is not optional boilerplate.
  • Source code (training, harness, evaluation scripts) is not released with this checkpoint. Only the merged weights, this card, the directive spec, and the JSON schemas / grammars for guided decoding are published.
  • On provenance, measured rather than assumed: asked "are you Qwen?", this model answers "No, I'm not Qwen. I'm Mark, Saanora's assistant" — it identifies itself correctly but does not volunteer its lineage, and in that phrasing it reads as a denial. Asked what model it is, it says it has no model name to report. Do not rely on the model to state its own provenance. Attribution for this release lives where the license actually requires it: in NOTICE, in this card, and in the base_model metadata above. If your product surfaces an identity answer, put the lineage in your system prompt (see Quickstart).
  • Training prompts include 252 human-written prompts from oasst2 (LAION e.V. and the Open Assistant contributors, Apache-2.0). Only the prompts were used; none of the oasst2 responses appear in the training data. Attribution is in NOTICE.

Built with Qwen.

Citation

@misc{mark1x9b2026,
  title  = {Mark 1x-9B (v1.1b): a 9B model that answers in interfaces},
  author = {Saathwik and {Saanora}},
  year   = {2026},
  note   = {LoRA r=32 adapter trained on Qwen/Qwen3.5-9B, merged to bf16},
  url    = {https://huggingface.co/Saanora/mark-1x-9b}
}

Please also cite Qwen3.5-9B, which this work builds on.

Downloads last month
305
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for Saanora/mark-1x-9b

Finetuned
Qwen/Qwen3.5-9B
Adapter
(785)
this model
Adapters
1 model

Evaluation results