Instructions to use StandardThinking/StandardOne-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use StandardThinking/StandardOne-3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="StandardThinking/StandardOne-3B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("StandardThinking/StandardOne-3B") model = AutoModelForMultimodalLM.from_pretrained("StandardThinking/StandardOne-3B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use StandardThinking/StandardOne-3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "StandardThinking/StandardOne-3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "StandardThinking/StandardOne-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/StandardThinking/StandardOne-3B
- SGLang
How to use StandardThinking/StandardOne-3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "StandardThinking/StandardOne-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "StandardThinking/StandardOne-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "StandardThinking/StandardOne-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "StandardThinking/StandardOne-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use StandardThinking/StandardOne-3B with Docker Model Runner:
docker model run hf.co/StandardThinking/StandardOne-3B
Download QUICKSTART.md from StandardThinking/StandardOne-3B: direct link, hf CLI and curl.
- Browser
- Download file 4.52 kB
-
https://huggingface.co/StandardThinking/StandardOne-3B/resolve/main/QUICKSTART.md
- Command line
-
hf download hf://StandardThinking/StandardOne-3B/QUICKSTART.md
-
curl -L -o QUICKSTART.md https://huggingface.co/StandardThinking/StandardOne-3B/resolve/main/QUICKSTART.md
Quickstart
Standard One 8B
- Clone the repository.
git clone https://huggingface.co/StandardThinking/StandardOne-8B
- Create a venv and install SGLang.
uv venv --python 3.12 .venv-sglang
uv pip install --python .venv-sglang/bin/python 'sglang==0.5.20'
- Start the engine.
CUDA_VISIBLE_DEVICES=0 SGLANG_VLM_CACHE_SIZE_MB=0 .venv-sglang/bin/python -m sglang.launch_server \
--model-path ./StandardOne-8B --served-model-name standard-one-8b \
--host 127.0.0.1 --port 30000 --tp-size 1 --model-impl sglang --dtype bfloat16 \
--context-length 32768 --max-running-requests 32 --mem-fraction-static 0.8 \
--chunked-prefill-size -1 --disable-radix-cache --mm-preprocess-cache-size-mb 0 \
--model-config-parser hf --load-format safetensors
- Create a venv and install the adapter.
uv venv --python 3.12 .venv-native
pip install -e ./StandardOne-8B/server[native-tokenizer]
- Start the adapter.
jev-adapter --engine-url http://127.0.0.1:30000 --model standard-one-8b --alias jev-latest \
--host 0.0.0.0 --port 30120 --max-concurrency 1 \
--tokenizer-model mistralai/Ministral-3-8B-Instruct-2512-BF16 \
--tokenizer-revision f6fae9795746f63c9be8344932f01275f3c63734 \
--prompt-wording served --label-scheme upper --default-temperature 1.65
- Check health and models.
curl -s http://127.0.0.1:30120/health
curl -s http://127.0.0.1:30120/v1/models
- Send a request.
curl -s http://127.0.0.1:30120/v1/systemone -X POST -H 'content-type: application/json' -d '{
"model": "jev-latest",
"state": "Policy: refunds require a receipt and purchase within 30 days. A customer bought 12 days ago but has no receipt. Issue a refund.",
"questions": {
"decision": {
"type": "noul",
"instructions": "Under the stated policy, is the requested action permitted? Treat unproved required conditions as not satisfied.",
"criteria": {"true": "Every required condition is established and no prohibition applies.", "false": "A condition is missing or a prohibition applies."}
}
}
}'
- Or run the packaged smoke script.
.venv-native/bin/python StandardOne-8B/server/examples/smoke.py --base-url http://127.0.0.1:30120 --request sample-request.json
- Run JevBench against the endpoint.
python -m jevbench.cli run --tasks jevbench-easy,jevbench-original,jevbench-hard \
--adapter typesafe --endpoint http://127.0.0.1:30120 --model jev-latest
Standard One 3B
These commands use the v2.2 3B checkpoint with the serving setup of the v2.2 measurements.
- Clone the repository.
git clone https://huggingface.co/StandardThinking/StandardOne-3B
- Start the engine (reuses the
.venv-sglangvenv from the 8B section).
.venv-sglang/bin/python -m sglang.launch_server \
--model-path ./StandardOne-3B --served-model-name standard-one-3b \
--host 127.0.0.1 --port 30000 --tp-size 1 --model-impl sglang --dtype bfloat16 \
--context-length 32768 --max-running-requests 32 --mem-fraction-static 0.8 \
--chunked-prefill-size -1 --disable-radix-cache --mm-preprocess-cache-size-mb 0 \
--model-config-parser hf --load-format safetensors
- Start the adapter (reuses
.venv-nativeand theStandardOne-8B/serverinstall from the 8B section).
.venv-native/bin/jev-adapter --engine-url http://127.0.0.1:30000 --model standard-one-3b --alias jev-latest \
--host 0.0.0.0 --port 30120 --max-concurrency 1 \
--tokenizer-model mistralai/Ministral-3-3B-Instruct-2512-BF16 \
--tokenizer-revision b6d637bef2393152b3da2b2fde72eecdee30557e \
--prompt-wording served --label-scheme upper --default-temperature 1.65
- Check health, then send a request or run JevBench exactly as in steps 6-9 above, with
"model": "jev-latest"unchanged (the adapter's alias, not the served checkpoint name).
Request format
POST /v1/systemone takes a state (the scenario, as text, and for supported task families an image
as a data URL) and a questions map. Each question has a type of noul (yes/no), choice (one of
several labeled options) or score (an ordinal scale), plus instructions and criteria describing
the labels. An optional options.temperature overrides the server's default softmax temperature for
that request. The response carries one native probability distribution per question; usage.output_tokens
is always 0, and every probability vector sums to 1 over exactly the caller's label set.