Spaces:
Running
Running
Update via HuggingBot
Browse files
README.md
CHANGED
|
@@ -1,3 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
# OpenJev — zero-GPU CPU deploy kit
|
| 2 |
|
| 3 |
Runs the [OpenJev](https://huggingface.co/openjev/openjev) decision API
|
|
@@ -15,41 +26,30 @@ output position — by pointing the **official, unmodified** `openjev-server` at
|
|
| 15 |
| latency | ~125 ms | seconds (see [Benchmarks](#benchmarks)) |
|
| 16 |
| surface | `/v1/systemone` + ops | identical (official server, unchanged) |
|
| 17 |
|
|
|
|
|
|
|
| 18 |
## Why this works (no code hacks)
|
| 19 |
|
| 20 |
-
`openjev-server`'s `vllm` backend is really *"any OpenAI-compatible server that
|
| 21 |
-
returns logprobs"*:
|
| 22 |
|
| 23 |
-
- it sends standard `/chat/completions` JSON (`logprobs`, `top_logprobs`,
|
| 24 |
-
|
| 25 |
-
- **exact readout** uses vLLM's `logprob_token_ids = [ids]` extension →
|
| 26 |
-
llama.cpp doesn't have it, so the backend's documented fallback kicks in:
|
| 27 |
-
top-K logprobs are matched **by label text** instead (see
|
| 28 |
-
`openjev_server/backends/vllm.py` → `_scores`, "old protocol").
|
| 29 |
|
| 30 |
-
That single fallback is the whole trick: the official code path runs unmodified,
|
| 31 |
-
and `--force` only relaxes the startup probe's letter-mass check if llama.cpp's
|
| 32 |
-
text-matched logprobs ever look borderline for your quant.
|
| 33 |
|
| 34 |
## What you get
|
| 35 |
|
| 36 |
- `POST /v1/systemone` — one request, any mix of `choice` / `noul` / `score`
|
| 37 |
- `POST /v1/prewarm`, `POST /v1/chat/completions` (passthrough)
|
| 38 |
-
- `GET /v1/version` · `GET /healthz` · `GET /readyz` ·
|
| 39 |
-
`GET /metrics` (Prometheus) · `GET /docs` (OpenAPI)
|
| 40 |
- optional bearer auth via `OPENJEV_TOKEN`
|
| 41 |
- Persistent model storage (HF Spaces) — the 16.2 GB download survives restarts
|
| 42 |
|
| 43 |
## Deploy on Hugging Face (zero GPU, free)
|
| 44 |
|
| 45 |
-
1. **Duplicate the runtime Space** (needs a free HF account):
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
**30 GB** ($0.50/mo, required for `/data`)
|
| 49 |
-
2. Wait for the build. The Docker image compiles llama.cpp CPU + installs
|
| 50 |
-
openjev-server and, on first boot, downloads the Q4 GGUF (~16.2 GB) into
|
| 51 |
-
`/data`. **Allow ~20–50 min on the first cold boot;** subsequent restarts
|
| 52 |
-
load straight from persistent storage.
|
| 53 |
3. Try the API. By default the Space is public:
|
| 54 |
|
| 55 |
```bash
|
|
@@ -83,12 +83,7 @@ curl -s "$BASE/v1/systemone" -H 'Content-Type: application/json' -d '{
|
|
| 83 |
}
|
| 84 |
```
|
| 85 |
|
| 86 |
-
> **Latency honesty right up front** — a 27B model on a shared free CPU is *not*
|
| 87 |
-
> 125 ms. Expect **~2–6 s per question** (single reader) on a 16-core Space with
|
| 88 |
-
> the Q4 quant. That's a real API — just not a high-QPS one. The reason this is
|
| 89 |
-
> still interesting: **a decision API only ever generates 1 token**, so the CPU
|
| 90 |
-
> is the *only* slow part, and it stays predictable. Raise `LLAMA_PARALLEL`,
|
| 91 |
-
> lower `LLAMA_CTX`, or use a different quant if your workload wants less of it.
|
| 92 |
|
| 93 |
## Deploy locally (or on any VPS)
|
| 94 |
|
|
@@ -124,8 +119,7 @@ docker run -d -p 7860:7860 -v openjev-data:/data -e MODEL_FILE=OpenJev-Q4_K_M.gg
|
|
| 124 |
|
| 125 |
## Benchmarks
|
| 126 |
|
| 127 |
-
Expected on a **free 16 vCPU CPU Space**, `OpenJev-Q4_K_M.gguf`
|
| 128 |
-
(27B, Qwen3-based → ~8-bit KV cache in the llama.cpp fork, 16.2 GB RAM):
|
| 129 |
|
| 130 |
| quant | file size | SPE (tokens/s, 8 threads) | RAM | 1-question latency* |
|
| 131 |
|---|---|---|---|---|
|
|
@@ -133,13 +127,9 @@ Expected on a **free 16 vCPU CPU Space**, `OpenJev-Q4_K_M.gguf`
|
|
| 133 |
| Q5_K_M | 18.8 GB | 1.5–3 | ~20 GB | ~3–8 s |
|
| 134 |
| Q6_K | 21.6 GB | 1.5–3 | ~23 GB | ~3–8 s |
|
| 135 |
|
| 136 |
-
\* single question ≈ one forward pass: `(prompt_tokens + 1) / SPE`. A 240-token
|
| 137 |
-
state ⇒ ~4 s @ 3 SPE. Latency is directly linear in prompt length.
|
| 138 |
|
| 139 |
-
**The decision-API twist:** because `max_tokens: 1`, a *pipeline* of many
|
| 140 |
-
questions on one state shares ~all of the prefill. For batch agents, batch the
|
| 141 |
-
questions into one `/v1/systemone` call — every extra answer after the first is
|
| 142 |
-
nearly free:
|
| 143 |
|
| 144 |
| questions in one call | est. per-call latency @ 3 SPE |
|
| 145 |
|---|---|
|
|
@@ -180,20 +170,21 @@ const r = await fetch(`${BASE}/v1/systemone`, {
|
|
| 180 |
const { answers } = await r.json();
|
| 181 |
```
|
| 182 |
|
| 183 |
-
##
|
| 184 |
|
| 185 |
-
|
| 186 |
-
|
| 187 |
-
|
| 188 |
-
|
| 189 |
-
|
| 190 |
-
|
| 191 |
-
|
| 192 |
-
``
|
|
|
|
|
|
|
| 193 |
|
| 194 |
## License & credit
|
| 195 |
|
| 196 |
-
- Model weights: **CC BY-NC 4.0** — OpenJev is an independent project, not
|
| 197 |
-
affiliated with TypeSafe (Jev is their product).
|
| 198 |
- `openjev-server`: Apache 2.0.
|
| 199 |
- This kit: Apache 2.0. llama.cpp: MIT.
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: OpenJev · Zero-GPU CPU Deploy Kit
|
| 3 |
+
emoji: ⚖️
|
| 4 |
+
colorFrom: green
|
| 5 |
+
colorTo: gray
|
| 6 |
+
sdk: static
|
| 7 |
+
pinned: true
|
| 8 |
+
license: apache-2.0
|
| 9 |
+
short_description: Zero-GPU CPU deploy kit for the OpenJev decision API
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
# OpenJev — zero-GPU CPU deploy kit
|
| 13 |
|
| 14 |
Runs the [OpenJev](https://huggingface.co/openjev/openjev) decision API
|
|
|
|
| 26 |
| latency | ~125 ms | seconds (see [Benchmarks](#benchmarks)) |
|
| 27 |
| surface | `/v1/systemone` + ops | identical (official server, unchanged) |
|
| 28 |
|
| 29 |
+
> 👀 The live index page is at **[https://broadfield-openjev-cpu-deploy.hf.space](https://broadfield-openjev-cpu-deploy.hf.space)** — this README is the machine-readable mirror. Every file below is fetchable raw: `https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/<path>`.
|
| 30 |
+
|
| 31 |
## Why this works (no code hacks)
|
| 32 |
|
| 33 |
+
`openjev-server`'s `vllm` backend is really *"any OpenAI-compatible server that returns logprobs"*:
|
|
|
|
| 34 |
|
| 35 |
+
- it sends standard `/chat/completions` JSON (`logprobs`, `top_logprobs`, `allowed_token_ids`, `max_tokens: 1`, `temperature: 0`);
|
| 36 |
+
- **exact readout** uses vLLM's `logprob_token_ids = [ids]` extension → llama.cpp doesn't have it, so the backend's documented fallback kicks in: top-K logprobs are matched **by label text** instead (see `openjev_server/backends/vllm.py` → `_scores`, "old protocol").
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
|
| 38 |
+
That single fallback is the whole trick: the official code path runs unmodified, and `--force` only relaxes the startup probe's letter-mass check if llama.cpp's text-matched logprobs ever look borderline for your quant.
|
|
|
|
|
|
|
| 39 |
|
| 40 |
## What you get
|
| 41 |
|
| 42 |
- `POST /v1/systemone` — one request, any mix of `choice` / `noul` / `score`
|
| 43 |
- `POST /v1/prewarm`, `POST /v1/chat/completions` (passthrough)
|
| 44 |
+
- `GET /v1/version` · `GET /healthz` · `GET /readyz` · `GET /metrics` (Prometheus) · `GET /docs` (OpenAPI)
|
|
|
|
| 45 |
- optional bearer auth via `OPENJEV_TOKEN`
|
| 46 |
- Persistent model storage (HF Spaces) — the 16.2 GB download survives restarts
|
| 47 |
|
| 48 |
## Deploy on Hugging Face (zero GPU, free)
|
| 49 |
|
| 50 |
+
1. **Duplicate the runtime Space** (needs a free HF account): <https://huggingface.co/spaces/broadfield/openjev-cpu-run?duplicate=true>
|
| 51 |
+
- SDK: **Docker** · hardware: **CPU basic (free)** · persistent storage: **30 GB** ($0.50/mo, required for `/data`)
|
| 52 |
+
2. Wait for the build. The Docker image compiles llama.cpp CPU + installs openjev-server and, on first boot, downloads the Q4 GGUF (~16.2 GB) into `/data`. **Allow ~20–50 min on the first cold boot;** subsequent restarts load straight from persistent storage.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 53 |
3. Try the API. By default the Space is public:
|
| 54 |
|
| 55 |
```bash
|
|
|
|
| 83 |
}
|
| 84 |
```
|
| 85 |
|
| 86 |
+
> **Latency honesty right up front** — a 27B model on a shared free CPU is *not* 125 ms. Expect **~2–6 s per question** (single reader) on a 16-core Space with the Q4 quant. That's a real API — just not a high-QPS one. The reason this is still interesting: **a decision API only ever generates 1 token**, so the CPU is the *only* slow part, and it stays predictable. Raise `LLAMA_PARALLEL`, lower `LLAMA_CTX`, or use a different quant if your workload wants less of it.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 87 |
|
| 88 |
## Deploy locally (or on any VPS)
|
| 89 |
|
|
|
|
| 119 |
|
| 120 |
## Benchmarks
|
| 121 |
|
| 122 |
+
Expected on a **free 16 vCPU CPU Space**, `OpenJev-Q4_K_M.gguf` (27B, Qwen3-based → ~8-bit KV cache in the llama.cpp fork, 16.2 GB RAM):
|
|
|
|
| 123 |
|
| 124 |
| quant | file size | SPE (tokens/s, 8 threads) | RAM | 1-question latency* |
|
| 125 |
|---|---|---|---|---|
|
|
|
|
| 127 |
| Q5_K_M | 18.8 GB | 1.5–3 | ~20 GB | ~3–8 s |
|
| 128 |
| Q6_K | 21.6 GB | 1.5–3 | ~23 GB | ~3–8 s |
|
| 129 |
|
| 130 |
+
\* single question ≈ one forward pass: `(prompt_tokens + 1) / SPE`. A 240-token state ⇒ ~4 s @ 3 SPE. Latency is directly linear in prompt length.
|
|
|
|
| 131 |
|
| 132 |
+
**The decision-API twist:** because `max_tokens: 1`, a *pipeline* of many questions on one state shares ~all of the prefill. For batch agents, batch the questions into one `/v1/systemone` call — every extra answer after the first is nearly free:
|
|
|
|
|
|
|
|
|
|
| 133 |
|
| 134 |
| questions in one call | est. per-call latency @ 3 SPE |
|
| 135 |
|---|---|
|
|
|
|
| 170 |
const { answers } = await r.json();
|
| 171 |
```
|
| 172 |
|
| 173 |
+
## Files in this kit (raw links)
|
| 174 |
|
| 175 |
+
| path | role | raw |
|
| 176 |
+
|---|---|---|
|
| 177 |
+
| `index.html` | live landing page with all instructions | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/index.html) |
|
| 178 |
+
| `Dockerfile` | multi-stage: llama.cpp CPU + official openjev-server | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/Dockerfile) |
|
| 179 |
+
| `entrypoint.sh` | GGUF → `/data`, llama-server, then `openjev serve` | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/entrypoint.sh) |
|
| 180 |
+
| `docker-compose.yml` | local/VPS one-liner | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/docker-compose.yml) |
|
| 181 |
+
| `bench.mjs` | latency/throughput smoke bench | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/bench.mjs) |
|
| 182 |
+
| `runtime-baseline.md` | the exact runtime Space spec | [raw](https://huggingface.co/spaces/broadfield/openjev-cpu-deploy/raw/main/runtime-baseline.md) |
|
| 183 |
+
|
| 184 |
+
Everything is also mirrored (cloneable) in the dataset repo [`broadfield/openjev-cpu-deploy-kit`](https://huggingface.co/datasets/broadfield/openjev-cpu-deploy-kit).
|
| 185 |
|
| 186 |
## License & credit
|
| 187 |
|
| 188 |
+
- Model weights: **CC BY-NC 4.0** — OpenJev is an independent project, not affiliated with TypeSafe (Jev is their product).
|
|
|
|
| 189 |
- `openjev-server`: Apache 2.0.
|
| 190 |
- This kit: Apache 2.0. llama.cpp: MIT.
|