Instructions to use Meanblock/JEV-CPU with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Meanblock/JEV-CPU with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("zero-shot-classification", model="Meanblock/JEV-CPU")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Meanblock/JEV-CPU", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 16,366 Bytes
759fa60 7845694 31ce914 7845694 af4853f 7845694 af4853f 7845694 8afa27f 7845694 8afa27f 7845694 be90e76 7845694 af4853f be90e76 7845694 af4853f 7845694 af4853f 0e01057 af4853f 9b0441f af4853f 7845694 9b0441f 7845694 af4853f 7845694 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 | ---
license: mit
base_model:
- Qwen/Qwen3-0.6B
tags:
- semantic-if
- decision
- semif
- cpu
- zero-shot-classification
pipeline_tag: zero-shot-classification
library_name: transformers
---
# JEV-CPU
<div align="center">
**Semantic ifs from open models β on a laptop CPU, no GPU.**
*A CPU port of [SemIf](https://github.com/TheoLeeCJ/SemIf) (formerly OpenJev), with a web UI.*
[](https://leesk212.github.io/JEV-CPU/)
[](https://github.com/leesk212/JEV-CPU)
[](https://huggingface.co/Meanblock/JEV-CPU)
[](./LICENSE)
π **[Read the full technical report β](https://leesk212.github.io/JEV-CPU/)** Β· [Run it locally](#quick-start) Β· [How it works](#how-a-decision-is-read-from-logits) Β· [Web UI](#web-ui)

*Live PoC β the state is **typed in**, criteria are added, and the decision is read from `Qwen3-0.6B`'s option logits in ~1 s (no text generated). One CPU engine across **eight domains**: support, content moderation, code-review triage, incident severity, email intent, compliance, loan/credit risk, and support prioritization.*
</div>
> **Independent project.** JEV-CPU is a thin CPU adaptation of [TheoLeeCJ/SemIf](https://github.com/TheoLeeCJ/SemIf). It is not affiliated with or endorsed by SemIf's author, TypeSafe, or Jev. Jev, TypeSafe, and other names and marks are the property of their respective owners. No infringement is intended.
Most agent decisions are small: *route this*, *retry that*, *does the evidence support X?* SemIf answers them by reading **typed option probabilities directly from a model** β no answer sentence, no JSON repair, no decoding loop. The upstream project targets a CUDA GPU holding a 4B BF16 model.
**JEV-CPU runs the exact same engine on a CPU**, with a small model and a browser UI, so you can try the pattern on any machine β no GPU, no waitlist.
---
## Why it runs on CPU
SemIf forces a GPU in exactly **one place** β `src/semif_phase1/core.py` β `load_causal_model()`:
```python
if not torch.cuda.is_available() or torch.cuda.device_count() != 1:
raise ValueError("Expose exactly one CUDA GPU ...")
...
dtype=torch.bfloat16, device_map={"": "cuda:0"} # β GPU pinned
```
Everything downstream is **device-agnostic**:
- `direct.py` / `shared.py` follow `device = next(model.parameters()).device`
- `torch.cuda.synchronize` is called **only** when `device.type == "cuda"` (a no-op on CPU)
- option-logit slot extraction and softmax are pure math
So JEV-CPU only swaps the loader (`semif_cpu.py`, CPU + `float32`) and reuses SemIf's **original, unmodified** scoring code.
```mermaid
flowchart LR
S[Unstructured state] --> M[Open model on CPU]
C[Runtime criteria] --> M
O[Typed options] --> M
M -- native option logits --> P[Probabilities]
```
- **Runtime-defined:** criteria and option descriptions arrive with the request.
- **Decision-native:** one forward pass reads declared option logits; no answer token is sampled.
- **No GPU:** loads `Qwen/Qwen3-0.6B` in `float32` (~2.4 GB) on CPU.
---
## How a decision is read from logits
This is the core idea, and it is why decisions are fast and cheap on a CPU: **the model never generates the answer β JEV-CPU reads it straight out of a single forward pass.** Below is the exact `direct` path (`src/semif_phase1/direct.py` + `core.py`), unchanged from upstream SemIf.
**1. Turn the decision into a letter-choice prompt.**
Each option is assigned an uppercase letter (`A`, `B`, `C`, β¦). The request becomes one chat turn:
```
system: Apply the supplied criterion to the supplied evidence. Choose exactly
one listed option. Respond with only its uppercase letter, with no
explanation or reasoning.
user: {"evidence": <state>,
"criterion": <question>,
"options": [{"letter": "A", "description": "..."},
{"letter": "B", "description": "..."}, ...]}
```
The chat template is applied with `add_generation_prompt=True` and `enable_thinking=False`, so the very next token the model would emit is the answer letter.
**2. Pin each option to exactly one token.**
For every letter, `_slot_ids()` checks that the letter encodes to a **single token** that round-trips (`decode(encode("A")) == "A"`), and that appending it to the prompt does not change the prompt's tokenization. This guarantees each option maps to one clean, comparable vocabulary slot β no multi-token drift, no whitespace merges.
**3. One forward pass β no decoding loop.**
The prompt is run through the model **once**. We take the logits at the final position only β the distribution over the *next* token:
```python
logits = model(**inputs, use_cache=False).logits[:, -1, :] # (vocab,)
```
No sampling, no `.generate()`, no answer sentence, no JSON to repair.
**4. Keep only the option slots, then softmax.**
From that full-vocabulary logit vector we gather just the option-letter token ids and softmax **over those slots alone**:
```python
selected = logits[slot_ids] # e.g. logits at tokens A, B, C
probs = softmax(selected) # conditional distribution over the options
winner = options[argmax(probs)]
```
The result is a probability per option, **conditional on the declared option set** β reported by JEV-CPU as `option_logits` + `probabilities`. Because it is one forward pass reading fixed positions, latency is dominated by prefill, not by generation length.
> **Why it is device-agnostic:** every step above is `device = next(model.parameters()).device`; the only CUDA-specific call in the whole path is `torch.cuda.synchronize()`, guarded by `if device.type == "cuda"`. On CPU it is simply skipped β so the *same* code runs unchanged.
### Reusing one state across many criteria (`shared` mode)
When every criterion judges the **same** state, `shared.py` prefills that state once into a native KV cache, replicates the cache across branches (`reorder_cache`), and evaluates all criteria's option positions in one batched forward using `logits_to_keep`. One expensive prefill, many cheap decisions β see upstream SemIf's speed tables. JEV-CPU inherits this path as-is (CPU just runs it without the CUDA sync).
---
## Quick start
Python 3.10+, ~3 GB RAM, no GPU:
```bash
python3 -m venv .venv && source .venv/bin/activate # Debian/Ubuntu: apt install python3-venv
pip install --index-url https://download.pytorch.org/whl/cpu torch
pip install transformers accelerate
# 1) CLI demo β prints typed option probabilities (semantic if)
python semif_cpu.py
# 2) Web UI β http://localhost:8080 (binds 0.0.0.0)
python server.py
```
The first run downloads `Qwen/Qwen3-0.6B` from Hugging Face and loads it on CPU (β 5β17 s). The model is loaded once and reused across requests.
---
## Web UI

`server.py` is a dependency-free (standard-library) web server on **port 8080**, laid out in three panes:
```
ββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββ
β β State / Evidence β β
β (data to judge) β β’ Results β
ββββββββββββββββββββββββββ€ per-option probability β
β β‘ Criteria β bars + chosen option β
β (add question+options)β β
ββββββββββββββββββββββββββ΄ββββββββββββββββββββββββββββ
```
- **β State** β the data to judge: a review, ticket, log line, or JSON blob.
- **β‘ Criteria** β add questions, each with two or more typed options, at runtime.
- **β’ Results** β for every criterion, the option probabilities the model read, and the winning choice.
### API
```
GET /api/health
POST /api/decide
{ "state": "...", "criteria": [ { "id", "question", "options": [ {"id","description"}, ... ] } ] }
```
---
## Verified results (CPU Β· Qwen3-0.6B Β· float32)
These are decisions from the demo GIF above β one CPU engine, eight domains:
| Domain | Criterion | Decision | Forward |
|---|---|---|---:|
| Customer support | Sentiment | **negative β 97.4%** β
| ~1.1 s |
| Customer support | Route to team | **billing β 100%** β
| ~1.0 s |
| Content moderation | Policy violation? | **violation β 99.3%** β
| ~1.1 s |
| Content moderation | Recommended action | **warn β 69.2%** β
| ~1.0 s |
| Code-review triage | Merge risk | **high β 99.8%** β
| ~1.2 s |
| Code-review triage | PR disposition | **block β 94.9%** β
| ~1.1 s |
| Incident / DevOps | Severity | **sev1 β 100%** β
| ~1.2 s |
| Incident / DevOps | Page on-call now? | **page_now β 100%** β
| ~1.1 s |
| Email intent | Primary intent | **sales β 100%** β
| ~1.2 s |
| Compliance gate | Change ticket required? | **required β 100%** β
| ~1.1 s |
| Loan / credit risk | Credit risk | **high β 100%** β
| ~1.2 s |
| Loan / credit risk | Recommended decision | **approve β 82.8%** β οΈ | ~1.1 s |
| Support prioritization | Priority | **p1 β 100%** β
| ~1.2 s |
- Model load β 5β17 s; each decision β **1 s** on CPU (no text is generated).
- Because SemIf reads option logits instead of decoding tokens, CPU latency stays low.
- β οΈ The loan **decision** row is a small-model slip: `Qwen3-0.6B` correctly flags *high risk* but still leans *approve* β an inconsistency that larger models resolve (see **[Scaling up](#scaling-up-the-brain-model)**).
### Per-domain demos (click to expand)
Each clip types the state in live, adds the criteria, and reads the decision from logits β on CPU.
<details>
<summary><b>π§ Customer support</b> β sentiment + team routing</summary>

</details>
<details>
<summary><b>π‘οΈ Content moderation</b> β policy violation + action</summary>

</details>
<details>
<summary><b>π Code-review triage</b> β merge risk + PR disposition</summary>

</details>
<details>
<summary><b>π¨ Incident / DevOps</b> β severity + page on-call</summary>

</details>
<details>
<summary><b>π§ Email intent</b> β intent classification</summary>

</details>
<details>
<summary><b>π Compliance gate</b> β change-ticket requirement</summary>

</details>
<details>
<summary><b>π³ Loan / credit risk</b> β risk + recommended decision</summary>

</details>
<details>
<summary><b>β±οΈ Support prioritization</b> β ticket priority</summary>

</details>
---
## Scaling up the brain model
The decisions above run on `Qwen/Qwen3-0.6B` β the **smallest** model on SemIf's ladder, chosen so it fits in CPU RAM. It is the accuracy floor, not the ceiling. From SemIf's own evaluation:
| Brain model | Size | Authored balanced accuracy | TypeSafe subset agreement |
|---|---:|---:|---:|
| **Qwen3-0.6B** (JEV-CPU default) | 0.6 B | 0.440 | 0.407 |
| MiniCPM5-2B | 2 B | 0.686 | 0.637 |
| **Qwen3.5-4B** | 4 B | **0.813** | **0.845** |
**Swapping the brain is a one-line change** β set `MODEL` in `semif_cpu.py`; the JEV-CPU engine and web UI are model-agnostic:
```python
MODEL = "openbmb/MiniCPM5-2B" # or "Qwen/Qwen3.5-4B"
```
That is exactly what fixes the slip in the loan demo: `Qwen3-0.6B` flags *high risk* correctly but still leans *approve*; a 2B/4B brain keeps the secondary decision consistent. The trade-off is resources β a 4B model needs `transformers`' native Qwen3.5 support and more RAM/compute than this 8 GB CPU box; a GPU (SemIf's target) makes it comfortable.
**Takeaway:** JEV-CPU shows the method runs anywhere; **accuracy scales with the model you point it at.** And a larger open-weight model **served on a GPU** relaxes the CPU latency wall too β lower latency, far larger inputs (toward the model's 40,960-token context), and many decisions per second via batching + shared-state reuse. That's the shape of a **production JEV** β see the report's [Outlook: GPU serving & a production JEV](https://leesk212.github.io/JEV-CPU/#outlook).
---
## Input token limits
Two different numbers matter β and the smaller one is **not** a limit of the small model:
| Limit | Value | What it is |
|---|---:|---|
| Model context (`Qwen3-0.6B`) | **40,960 tokens** | The model's architectural window β large even at 0.6 B; context length comes from RoPE, independent of parameter count. |
| JEV-CPU / SemIf default cap | **4,096 tokens / decision** | A safety guard (`max_tokens`); over-long prompts raise instead of being silently truncated. Configurable. |
| Practical max on this 8 GB CPU box | **β 7,700 tokens (~117 s)** | Where CPU **prefill latency** becomes the ceiling. Beyond this a single decision crosses **~120 s**, which we treat as impractical β not a memory or model limit. |
**Measured on this 8 GB CPU box** (no GPU), one decision, growing input:
| Input tokens | Forward (CPU) | Peak RAM |
|---:|---:|---:|
| 254 | 2.5 s | ~3.5 GB |
| 731 | 5.6 s | ~3.5 GB |
| 1,363 | 11.9 s | ~3.5 GB |
| 2,629 | 26.0 s | ~3.5 GB |
| 5,165 | 66.1 s | ~3.5 GB |
| 7,697 | 117.4 s | ~3.5 GB |
RAM stayed **flat at ~3.5 GB** even at 7,697 tokens β well past the 4,096 default and with no OOM β so on this box the ceiling is **prefill latency (β quadratic)**, not memory or the model. We cap the **practical input at β 7,700 tokens (~117 s)**: past that, one decision exceeds **~120 s**, which is no longer useful on CPU, so larger inputs are treated as unsupported here. Interactive ~1 s decisions want short states (β² ~300 tokens). A GPU removes this latency wall entirely β where much larger inputs (up to the model's 40,960) become usable again.
Raising the cap is a parameter, not a rebuild: pass a larger `max_tokens` to `direct_score(...)` (or `--max-tokens` in upstream SemIf's CLI). The model accepts input up to 40,960 tokens; on **CPU** the real constraint is prefill time β latency grows with length, so short states keep decisions near ~1 s.
---
## Input
```json
{
"id": "route-1",
"state": "Customer cannot access an account after a password reset.",
"question": "Which queue should handle this request?",
"options": [
{"id": "access", "description": "Account access support."},
{"id": "billing", "description": "Billing support."}
]
}
```
Returned probabilities are conditional on the supplied options β calibrate them on your workload. `state` may also be a nonempty JSON object or array.
---
## What's in this repo
| Path | Description |
|---|---|
| `semif_cpu.py` | CPU/`float32` loader shim + `direct.score()` example. Auto-detects SemIf source (`SEMIF_DIR` env β repo `./src` β `/tmp/SemIf`). |
| `server.py` | Standard-library web server (port 8080) + three-pane UI. |
| `src/semif_phase1/` | Upstream SemIf engine, **unchanged**. |
| `README.SemIf-upstream.md` | Original SemIf README. |
---
## Credits & license
- Engine: **[TheoLeeCJ/SemIf](https://github.com/TheoLeeCJ/SemIf)** β *"Semantic ifs from open models."* Browser demo: <https://openjev.com/>
- Interface concept: TypeSafe's *Jev* pattern.
- Baseline model: [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B).
JEV-CPU adds only a CPU loader shim and a web UI on top of SemIf; the scoring logic is unchanged. Upstream models retain their licenses; see [`THIRD_PARTY.md`](./THIRD_PARTY.md). Project code is released under the [MIT License](./LICENSE).
|