broadfield commited on
Commit
a4cfbd5
·
verified ·
1 Parent(s): f321f11

Update via HuggingBot

Browse files
Files changed (1) hide show
  1. README.md +37 -46
README.md CHANGED
@@ -1,3 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
1
  # OpenJev — zero-GPU CPU deploy kit
2
 
3
  Runs the [OpenJev](https://huggingface.co/openjev/openjev) decision API
@@ -15,41 +26,30 @@ output position — by pointing the **official, unmodified** `openjev-server` at
15
  | latency | ~125 ms | seconds (see [Benchmarks](#benchmarks)) |
16
  | surface | `/v1/systemone` + ops | identical (official server, unchanged) |
17
 
 
 
18
  ## Why this works (no code hacks)
19
 
20
- `openjev-server`'s `vllm` backend is really *"any OpenAI-compatible server that
21
- returns logprobs"*:
22
 
23
- - it sends standard `/chat/completions` JSON (`logprobs`, `top_logprobs`,
24
- `allowed_token_ids`, `max_tokens: 1`, `temperature: 0`);
25
- - **exact readout** uses vLLM's `logprob_token_ids = [ids]` extension →
26
- llama.cpp doesn't have it, so the backend's documented fallback kicks in:
27
- top-K logprobs are matched **by label text** instead (see
28
- `openjev_server/backends/vllm.py` → `_scores`, "old protocol").
29
 
30
- That single fallback is the whole trick: the official code path runs unmodified,
31
- and `--force` only relaxes the startup probe's letter-mass check if llama.cpp's
32
- text-matched logprobs ever look borderline for your quant.
33
 
34
  ## What you get
35
 
36
  - `POST /v1/systemone` — one request, any mix of `choice` / `noul` / `score`
37
  - `POST /v1/prewarm`, `POST /v1/chat/completions` (passthrough)
38
- - `GET /v1/version` · `GET /healthz` · `GET /readyz` ·
39
- `GET /metrics` (Prometheus) · `GET /docs` (OpenAPI)
40
  - optional bearer auth via `OPENJEV_TOKEN`
41
  - Persistent model storage (HF Spaces) — the 16.2 GB download survives restarts
42
 
43
  ## Deploy on Hugging Face (zero GPU, free)
44
 
45
- 1. **Duplicate the runtime Space** (needs a free HF account):
46
- <https://huggingface.co/spaces/broadfield/openjev-cpu-run?duplicate=true>
47
- - SDK: **Docker** · hardware: **CPU basic (free)** · persistent storage:
48
- **30 GB** ($0.50/mo, required for `/data`)
49
- 2. Wait for the build. The Docker image compiles llama.cpp CPU + installs
50
- openjev-server and, on first boot, downloads the Q4 GGUF (~16.2 GB) into
51
- `/data`. **Allow ~20–50 min on the first cold boot;** subsequent restarts
52
- load straight from persistent storage.
53
  3. Try the API. By default the Space is public:
54
 
55
  ```bash
@@ -83,12 +83,7 @@ curl -s "$BASE/v1/systemone" -H 'Content-Type: application/json' -d '{
83
  }
84
  ```
85
 
86
- > **Latency honesty right up front** — a 27B model on a shared free CPU is *not*
87
- > 125 ms. Expect **~2–6 s per question** (single reader) on a 16-core Space with
88
- > the Q4 quant. That's a real API — just not a high-QPS one. The reason this is
89
- > still interesting: **a decision API only ever generates 1 token**, so the CPU
90
- > is the *only* slow part, and it stays predictable. Raise `LLAMA_PARALLEL`,
91
- > lower `LLAMA_CTX`, or use a different quant if your workload wants less of it.
92
 
93
  ## Deploy locally (or on any VPS)
94
 
@@ -124,8 +119,7 @@ docker run -d -p 7860:7860 -v openjev-data:/data -e MODEL_FILE=OpenJev-Q4_K_M.gg
124
 
125
  ## Benchmarks
126
 
127
- Expected on a **free 16 vCPU CPU Space**, `OpenJev-Q4_K_M.gguf`
128
- (27B, Qwen3-based → ~8-bit KV cache in the llama.cpp fork, 16.2 GB RAM):
129
 
130
  | quant | file size | SPE (tokens/s, 8 threads) | RAM | 1-question latency* |
131
  |---|---|---|---|---|
@@ -133,13 +127,9 @@ Expected on a **free 16 vCPU CPU Space**, `OpenJev-Q4_K_M.gguf`
133
  | Q5_K_M | 18.8 GB | 1.5–3 | ~20 GB | ~3–8 s |
134
  | Q6_K | 21.6 GB | 1.5–3 | ~23 GB | ~3–8 s |
135
 
136
- \* single question ≈ one forward pass: `(prompt_tokens + 1) / SPE`. A 240-token
137
- state ⇒ ~4 s @ 3 SPE. Latency is directly linear in prompt length.
138
 
139
- **The decision-API twist:** because `max_tokens: 1`, a *pipeline* of many
140
- questions on one state shares ~all of the prefill. For batch agents, batch the
141
- questions into one `/v1/systemone` call — every extra answer after the first is
142
- nearly free:
143
 
144
  | questions in one call | est. per-call latency @ 3 SPE |
145
  |---|---|
@@ -180,20 +170,21 @@ const r = await fetch(`${BASE}/v1/systemone`, {
180
  const { answers } = await r.json();
181
  ```
182
 
183
- ## Layout
184
 
185
- ```
186
- Dockerfile # multi-stage: llama.cpp CPU + official openjev-server
187
- entrypoint.sh # GGUF → /data, llama-server, then openjev serve
188
- docker-compose.yml # local/VPS one-liner
189
- bench.mjs # throughput + latency smoke bench against a running endpoint
190
- runtime-baseline.md # the exact Docker Space spec we recommend
191
- README.md # you are here
192
- ```
 
 
193
 
194
  ## License & credit
195
 
196
- - Model weights: **CC BY-NC 4.0** — OpenJev is an independent project, not
197
- affiliated with TypeSafe (Jev is their product).
198
  - `openjev-server`: Apache 2.0.
199
  - This kit: Apache 2.0. llama.cpp: MIT.
 
1
+ ---
2
+ title: OpenJev · Zero-GPU CPU Deploy Kit
3
+ emoji: ⚖️
4
+ colorFrom: green
5
+ colorTo: gray
6
+ sdk: static
7
+ pinned: true
8
+ license: apache-2.0
9
+ short_description: Zero-GPU CPU deploy kit for the OpenJev decision API
10
+ ---
11
+
12
  # OpenJev — zero-GPU CPU deploy kit
13
 
14
  Runs the [OpenJev](https://huggingface.co/openjev/openjev) decision API
 
26
  | latency | ~125 ms | seconds (see [Benchmarks](#benchmarks)) |
27
  | surface | `/v1/systemone` + ops | identical (official server, unchanged) |
28
 
29
+ > 👀 The live index page is at **[https://broadfield-openjev-cpu-deploy.hf.space](https://broadfield-openjev-cpu-deploy.hf.space)** — this README is the machine-readable mirror. Every file below is fetchable raw: `https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/<path>`.
30
+
31
  ## Why this works (no code hacks)
32
 
33
+ `openjev-server`'s `vllm` backend is really *"any OpenAI-compatible server that returns logprobs"*:
 
34
 
35
+ - it sends standard `/chat/completions` JSON (`logprobs`, `top_logprobs`, `allowed_token_ids`, `max_tokens: 1`, `temperature: 0`);
36
+ - **exact readout** uses vLLM's `logprob_token_ids = [ids]` extension → llama.cpp doesn't have it, so the backend's documented fallback kicks in: top-K logprobs are matched **by label text** instead (see `openjev_server/backends/vllm.py` → `_scores`, "old protocol").
 
 
 
 
37
 
38
+ That single fallback is the whole trick: the official code path runs unmodified, and `--force` only relaxes the startup probe's letter-mass check if llama.cpp's text-matched logprobs ever look borderline for your quant.
 
 
39
 
40
  ## What you get
41
 
42
  - `POST /v1/systemone` — one request, any mix of `choice` / `noul` / `score`
43
  - `POST /v1/prewarm`, `POST /v1/chat/completions` (passthrough)
44
+ - `GET /v1/version` · `GET /healthz` · `GET /readyz` · `GET /metrics` (Prometheus) · `GET /docs` (OpenAPI)
 
45
  - optional bearer auth via `OPENJEV_TOKEN`
46
  - Persistent model storage (HF Spaces) — the 16.2 GB download survives restarts
47
 
48
  ## Deploy on Hugging Face (zero GPU, free)
49
 
50
+ 1. **Duplicate the runtime Space** (needs a free HF account): <https://huggingface.co/spaces/broadfield/openjev-cpu-run?duplicate=true>
51
+ - SDK: **Docker** · hardware: **CPU basic (free)** · persistent storage: **30 GB** ($0.50/mo, required for `/data`)
52
+ 2. Wait for the build. The Docker image compiles llama.cpp CPU + installs openjev-server and, on first boot, downloads the Q4 GGUF (~16.2 GB) into `/data`. **Allow ~20–50 min on the first cold boot;** subsequent restarts load straight from persistent storage.
 
 
 
 
 
53
  3. Try the API. By default the Space is public:
54
 
55
  ```bash
 
83
  }
84
  ```
85
 
86
+ > **Latency honesty right up front** — a 27B model on a shared free CPU is *not* 125 ms. Expect **~2–6 s per question** (single reader) on a 16-core Space with the Q4 quant. That's a real API — just not a high-QPS one. The reason this is still interesting: **a decision API only ever generates 1 token**, so the CPU is the *only* slow part, and it stays predictable. Raise `LLAMA_PARALLEL`, lower `LLAMA_CTX`, or use a different quant if your workload wants less of it.
 
 
 
 
 
87
 
88
  ## Deploy locally (or on any VPS)
89
 
 
119
 
120
  ## Benchmarks
121
 
122
+ Expected on a **free 16 vCPU CPU Space**, `OpenJev-Q4_K_M.gguf` (27B, Qwen3-based → ~8-bit KV cache in the llama.cpp fork, 16.2 GB RAM):
 
123
 
124
  | quant | file size | SPE (tokens/s, 8 threads) | RAM | 1-question latency* |
125
  |---|---|---|---|---|
 
127
  | Q5_K_M | 18.8 GB | 1.5–3 | ~20 GB | ~3–8 s |
128
  | Q6_K | 21.6 GB | 1.5–3 | ~23 GB | ~3–8 s |
129
 
130
+ \* single question ≈ one forward pass: `(prompt_tokens + 1) / SPE`. A 240-token state ⇒ ~4 s @ 3 SPE. Latency is directly linear in prompt length.
 
131
 
132
+ **The decision-API twist:** because `max_tokens: 1`, a *pipeline* of many questions on one state shares ~all of the prefill. For batch agents, batch the questions into one `/v1/systemone` call — every extra answer after the first is nearly free:
 
 
 
133
 
134
  | questions in one call | est. per-call latency @ 3 SPE |
135
  |---|---|
 
170
  const { answers } = await r.json();
171
  ```
172
 
173
+ ## Files in this kit (raw links)
174
 
175
+ | path | role | raw |
176
+ |---|---|---|
177
+ | `index.html` | live landing page with all instructions | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/index.html) |
178
+ | `Dockerfile` | multi-stage: llama.cpp CPU + official openjev-server | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/Dockerfile) |
179
+ | `entrypoint.sh` | GGUF → `/data`, llama-server, then `openjev serve` | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/entrypoint.sh) |
180
+ | `docker-compose.yml` | local/VPS one-liner | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/docker-compose.yml) |
181
+ | `bench.mjs` | latency/throughput smoke bench | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/bench.mjs) |
182
+ | `runtime-baseline.md` | the exact runtime Space spec | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/runtime-baseline.md) |
183
+
184
+ Everything is also mirrored (cloneable) in the dataset repo [`broadfield/openjev-cpu-deploy-kit`](https://huggingface.co/datasets/broadfield/openjev-cpu-deploy-kit).
185
 
186
  ## License & credit
187
 
188
+ - Model weights: **CC BY-NC 4.0** — OpenJev is an independent project, not affiliated with TypeSafe (Jev is their product).
 
189
  - `openjev-server`: Apache 2.0.
190
  - This kit: Apache 2.0. llama.cpp: MIT.