StandardOne-3B-FP8 / README.md
MyeongHoJeong's picture
Update model card
094ab54 verified
|
Raw History Blame Contribute Delete
2.39 kB
---
base_model: StandardThinking/StandardOne-3B
library_name: transformers
license: apache-2.0
pipeline_tag: text-generation
tags:
- mistral3
- decision-model
- typed-decisions
- jev
- jevbench
- fp8
- compressed-tensors
- vision-language
---
# StandardOne-3B-FP8
**Version:** v2.2
StandardOne-3B-FP8 is an FP8 (compressed-tensors, float8_e4m3 weights, dynamic per-token activations)
quantization of the released StandardOne-3B decision model. The language-model linear
projections (q/k/v/o, gate/up/down) are quantized per-channel FP8 E4M3 with dynamic FP8
activations (llm-compressor's data-free `FP8_DYNAMIC` recipe, no calibration data required);
the vision tower, multi-modal projector, embeddings and lm_head are left unquantized in BF16.
It was produced from StandardOne-3B v2.2 (tag `v2.2`) on 2026-10-04 using
llm-compressor 0.14.0 (torch 2.14.0, transformers 5.17.0,
compressed-tensors 0.19.0); results below.
## Changes in v2.2
Rebuilt from StandardOne-3B v2.2 with the same recipe (`recipe.yaml` unchanged). The v2.1 to v2.2 score changes are listed in the [StandardOne-3B card](https://huggingface.co/StandardThinking/StandardOne-3B#changes-in-v22).
## Validation
Both precisions were served the same way through SGLang 0.5.20 and `jev-adapter` (`served` wording, one option order, default temperature 1.65), measured 2026-10-04. Accuracy is the most probable answer and does not depend on temperature.
| Suite | BF16 (StandardOne-3B v2.2) | FP8 (this repository) | Change (points) |
|---|---:|---:|---:|
| many-option questions, 53–151 options (18,000) | 79.97 % | 79.57 % | −0.40 |
| the same question set, at most 26 options (750) | 85.47 % | 84.67 % | −0.80 |
| long-document questions (150) | 38.00 % | 44.00 % | +6.00 |
| held-out decision set (600) | 82.00 % | 80.50 % | −1.50 |
| hard proxy (600) | 44.67 % | 45.17 % | +0.50 |
| realistic transfer set (600) | 88.17 % | 87.33 % | −0.84 |
| JevBench public easy (48) | 100.00 % | 97.92 % | −2.08 |
| JevBench public standard (72) | 93.06 % | 91.67 % | −1.39 |
| JevBench public hard (111) | 45.95 % | 44.14 % | −1.81 |
Across 28,406 validation questions with recorded probabilities, FP8 and BF16 gave the same answer for 95.22 %. No separate temperature was fitted for this build; use the serving settings of [StandardOne-3B](https://huggingface.co/StandardThinking/StandardOne-3B).