Text Classification
Transformers
Safetensors
Thai
English
openthai_systemone
feature-extraction
system-one
decision-model
thai
qwen3.5
quantized
compressed-tensors
llm-compressor
custom_code
Instructions to use iapp/OpenThai-SystemOne-FP8-Dynamic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use iapp/OpenThai-SystemOne-FP8-Dynamic with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="iapp/OpenThai-SystemOne-FP8-Dynamic", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("iapp/OpenThai-SystemOne-FP8-Dynamic", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from iapp/OpenThai-SystemOne-FP8-Dynamic: direct link, hf CLI and curl.
- Browser
- Download file 4.25 kB
-
https://huggingface.co/iapp/OpenThai-SystemOne-FP8-Dynamic/resolve/main/README.md
- Command line
-
hf download hf://iapp/OpenThai-SystemOne-FP8-Dynamic/README.md
-
curl -L -o README.md https://huggingface.co/iapp/OpenThai-SystemOne-FP8-Dynamic/resolve/main/README.md
4.25 kB
| license: apache-2.0 | |
| language: [th, en] | |
| base_model: iapp/OpenThai-SystemOne | |
| base_model_relation: quantized | |
| library_name: transformers | |
| pipeline_tag: text-classification | |
| tags: [system-one, decision-model, thai, qwen3.5, quantized, compressed-tensors, llm-compressor] | |
| # OpenThai-SystemOne — FP8-Dynamic | |
| [OpenThai-SystemOne](https://huggingface.co/iapp/OpenThai-SystemOne) is an open Thai + English **System One decision model**: one forward | |
| pass answers typed questions (`choice` over up to 255 options, ordinal `score`, yes/no `noul`) about a text / JSON state with | |
| calibrated probabilities, no text generation. It is a Qwen3.5-0.8B text tower (Thai continued pre-training) plus a 256-slot | |
| decision head. This repo is a quantization of **v0.3** (commit `f3709948`). | |
| **What is quantized:** the Linear layers of the tower. The token embeddings, the 256-slot decision head and the per-type temperatures stay in bf16. Quantization therefore only perturbs the hidden state the head reads. | |
| **Format:** FP8 W8A8 (dynamic per-token activations). compressed-tensors (llm-compressor) checkpoint. Loaded through transformers the weights are decompressed to bf16 at load time (same speed as bf16, smaller download); native FP8 / INT8 / FP4 kernels need a runtime with this architecture (the decision head is custom, so vLLM does not serve it out of the box). | |
| Size: 1009 MB (bf16 original: 1,509 MB). | |
| ## Usage | |
| ```bash | |
| pip install torch transformers safetensors pydantic && pip install compressed-tensors | |
| ``` | |
| ```python | |
| from transformers import AutoModel, AutoTokenizer # trust_remote_code files are in this repo | |
| model = AutoModel.from_pretrained("iapp/OpenThai-SystemOne-FP8-Dynamic", trust_remote_code=True) | |
| # or with the pip package (git+https://github.com/iapp-technology/openthai-systemone): | |
| from openthai_systemone import SystemOneClient | |
| c = SystemOneClient("iapp/OpenThai-SystemOne-FP8-Dynamic") | |
| r = c.system_one("ร้านนี้อาหารอร่อยมาก แต่รอนานเกือบชั่วโมง", {"sentiment": {"type": "choice", "instructions": "ความรู้สึก", | |
| "criteria": {"บวก": None, "ลบ": None, "กลาง": None}}}) | |
| ``` | |
| ## Accuracy vs the bf16 original (same records, single option order, first 800 per set) | |
| Macro: public 74.1 (original 74.3), Thai 80.1 (original 80.1). | |
| | subset | bf16 original | this | Δ | | |
| |---|---|---|---| | |
| | **public 13-subset bench** | | | | | |
| | aegis2 (noul) | 83.2 | 83.2 | +0.0 | | |
| | boolq (noul) | 79.7 | 80.0 | +0.3 | | |
| | civil_comments (noul) | 79.0 | 79.0 | +0.0 | | |
| | helpsteer2 (score) | 41.6 | 41.6 | +0.0 | | |
| | massive-de-DE (choice) | 88.3 | 87.7 | -0.6 | | |
| | massive-en-US (choice) | 88.3 | 88.0 | -0.3 | | |
| | multinli (choice) | 89.0 | 89.3 | +0.3 | | |
| | paws (noul) | 94.0 | 93.2 | -0.8 | | |
| | pubmedqa (choice) | 64.0 | 64.8 | +0.8 | | |
| | squad2 (noul) | 89.3 | 88.6 | -0.7 | | |
| | summeval-consistency (score) | 75.0 | 75.7 | +0.7 | | |
| | summeval-relevance (score) | 21.7 | 20.8 | -0.8 | | |
| | vitaminc-dev (choice) | 72.5 | 72.0 | -0.5 | | |
| | *macro, public 13-subset bench* | *74.3* | *74.1* | *-0.1* | | |
| | **Thai held-out / eval sets** | | | | | |
| | banking77 (choice) | 59.1 | 59.4 | +0.2 | | |
| | contrastive_th (choice) | 80.7 | 80.4 | -0.3 | | |
| | contrastive_th (noul) | 83.5 | 83.1 | -0.4 | | |
| | contrastive_th (score) | 78.6 | 78.6 | +0.0 | | |
| | massive_th (choice) | 90.6 | 90.5 | -0.1 | | |
| | prachathai (choice) | 98.3 | 98.5 | +0.2 | | |
| | prachathai (noul) | 93.4 | 93.4 | -0.1 | | |
| | sib200_th (choice) | 77.9 | 77.0 | -1.0 | | |
| | wisesight (choice) | 48.9 | 49.6 | +0.8 | | |
| | wongnai (score) | 64.5 | 65.0 | +0.5 | | |
| | xlam_tools (choice) | 99.4 | 99.4 | +0.0 | | |
| | xnli_th (choice) | 79.8 | 79.6 | -0.1 | | |
| | xnli_th (noul) | 86.8 | 86.6 | -0.1 | | |
| | *macro, Thai held-out / eval sets* | *80.1* | *80.1* | *-0.0* | | |
| ## Notes | |
| - Scores are single-option-order accuracy on the first 800 records of each set (`scripts/06_eval.py --limit 800`), the same | |
| records for the original and the quantization. `score` subsets report exact level accuracy. | |
| - Base model, data, training and the full benchmark tables: [iapp/OpenThai-SystemOne](https://huggingface.co/iapp/OpenThai-SystemOne). | |
| - License Apache-2.0 (same as the base). Built by iApp Technology / OpenThaiGPT. | |