111

A chat model fine-tuned from Qwen3.5-9B. It takes a structured prompt and returns a strict JSON response.

Serving

OpenAI-compatible chat model (served with an inference engine such as SGLang / vLLM):

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
    model="111",
    messages=[
        {"role": "system", "content": system_prompt},
        {"role": "user", "content": user_prompt},
    ],
    temperature=0.0,
    max_tokens=8192,
)
print(resp.choices[0].message.content)

Output format

Strict JSON, e.g.:

{"thought_process": "...", "valid_steps": [1, 2, 5, 8]}

Details

  • Base: Qwen3.5-9B (Qwen3_5ForConditionalGeneration, 32 layers, hidden size 4096).
  • Precision: bfloat16.
  • Format: safetensors (4 shards) + HF config.json / tokenizer.

License

Inherits the base-model license. Set the correct license before publishing.

Downloads last month
37
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support