Text Generation
Transformers
Safetensors
mistral3
image-text-to-text
decision-model
typed-decisions
jev
jevbench
calibration
decode-free
multilingual
vision-language
conversational
Instructions to use StandardThinking/StandardOne-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use StandardThinking/StandardOne-3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="StandardThinking/StandardOne-3B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("StandardThinking/StandardOne-3B") model = AutoModelForMultimodalLM.from_pretrained("StandardThinking/StandardOne-3B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use StandardThinking/StandardOne-3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "StandardThinking/StandardOne-3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "StandardThinking/StandardOne-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/StandardThinking/StandardOne-3B
- SGLang
How to use StandardThinking/StandardOne-3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "StandardThinking/StandardOne-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "StandardThinking/StandardOne-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "StandardThinking/StandardOne-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "StandardThinking/StandardOne-3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use StandardThinking/StandardOne-3B with Docker Model Runner:
docker model run hf.co/StandardThinking/StandardOne-3B
File size: 3,798 Bytes
ea84b46 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 | # KEV metric parity validation
On 2026-09-21, the adapter's scalar quality metrics were compared numerically
with the original KEV implementations at commit
`4f8110a3f8620cc3a182ae9a708e4398492c4b1a`.
All common scalar metrics matched within `1e-12` absolute error. The largest
observed difference was `8.881784197001252e-16` (score MAE). Complete contrastive
pair metrics matched exactly.
This validates the metric calculations. It does not validate H200 execution,
inference probabilities, model accuracy, or serving latency.
## Method
The optional [verify_metrics.py](verify_metrics.py) script reads these original
files from a local KEV checkout and verifies their full SHA256 before execution:
| File | Extracted functions | SHA256 |
|---|---|---|
| `kev/evaluate.py` | `ece` | `1b20e3f9edf417aa8dae924b1526e52f74b710cadf7213c5ec68334f6e7f8fe1` |
| `kev/benchmark.py` | `coverage_at_error`, `metrics` | `4192ec3b26b065452f84bde38a091e6854a28fe185d2e0f39a0c91b2efe69df7` |
| `kev/contrastive.py` | `paired_flip` | `cbb979aa5d40265ad0e64695f94b281d91751fa405811ddfa8212fede111edcf` |
Python AST extraction keeps only those function definitions. KEV's package,
PyTorch, Transformers and model-loading code are not imported. The audit uses
NumPy in its own optional environment; NumPy is not an adapter dependency.
The original successful audit used NumPy `2.3.5`.
Input generation uses `numpy.random.default_rng(483)` for 100 batches of 50
rows, totaling 5,000 rows. Rows mix choice, boolean and ordinal score questions,
with 2–10 options. Probabilities come from Dirichlet distributions. The first
10 rows in each batch additionally cover binary probability endpoints and
decimal boundaries, including zero and one. Both implementations receive the
same rows in the same order, including confidence ties.
Compared metrics include accuracy, NLL, multiclass Brier score, ten-bin ECE,
mean confidence, confidence bias, confident errors, coverage/accuracy at 0.9,
coverage at 1%/5% empirical error, score MAE and ranked probability score.
Ten complete two-sibling pairs additionally check relevant and invariant pair
metrics. Five pairs have changed gold labels and five have unchanged labels.
## Reproduce
Use a Python environment that already has NumPy installed:
```bash
git clone https://github.com/jaredpalmer/kev.git /tmp/kev-reference
git -C /tmp/kev-reference checkout 4f8110a3f8620cc3a182ae9a708e4398492c4b1a
python benchmarks/verify_metrics.py --kev-root /tmp/kev-reference
```
The JSON output includes the reference hashes, current adapter metric-code hash,
NumPy version, sample counts, per-metric maximum differences and pair results.
A source mismatch or numerical discrepancy exits unsuccessfully. Reference
files are checked even if the local checkout has uncommitted changes.
## Scope differences
- Accuracy headlines use clean question rows, not all submitted records.
Source `unknowable` is excluded from accuracy and scored for confidence.
- Incomplete pairs are counted explicitly for smoke subsets; KEV's original
pair function rejects incomplete pairs.
- The adapter omits meaningless unknowable accuracy from grouped reports;
KEV's original code includes it in some diagnostic subreports.
- This audit compares common scalar metrics and complete pairs. It does not
certify every report field, input conversion, HTTP behavior, or SemIf's
separate family-balanced aggregation.
Original definitions: [KEV benchmark.py](https://github.com/jaredpalmer/kev/blob/4f8110a3f8620cc3a182ae9a708e4398492c4b1a/kev/benchmark.py),
[evaluate.py](https://github.com/jaredpalmer/kev/blob/4f8110a3f8620cc3a182ae9a708e4398492c4b1a/kev/evaluate.py),
[contrastive.py](https://github.com/jaredpalmer/kev/blob/4f8110a3f8620cc3a182ae9a708e4398492c4b1a/kev/contrastive.py).
|