EchoCache / deploy.md
NagaYu's picture
EchoCache v1.0.0: docs, benchmark results and the full application source
46985e0 verified
|
Raw History Blame Contribute Delete
14 kB
# Deploying EchoCache on a free Hugging Face Space
> 日本語版: [`deploy.ja.md`](deploy.ja.md)
About 10 minutes, copy-paste throughout. You need a browser, `git` and `python3`.
No paid plan, no GPU, no model download.
---
## 1. Create the Space
1. Open <https://huggingface.co/new-space>
2. Fill in:
| Field | Value |
|---|---|
| **Owner** | your account |
| **Space name** | `EchoCache` (referred to below as `<user>/EchoCache`) |
| **License** | `apache-2.0` |
| **SDK** | **Gradio** |
| **Space hardware** | **CPU basic · 2 vCPU · 16 GB · FREE** |
| **Visibility** | Public or Private — both work |
3. Press **Create Space**. If a Gradio template picker appears, choose "Blank"; every file gets replaced anyway.
---
## 2. Push the files
Log in once with a **write** token from <https://huggingface.co/settings/tokens>:
```bash
hf auth login # older installs: huggingface-cli login
```
Then run this as-is, replacing `<user>` with your account name:
```bash
USER=<user>; SPACE=EchoCache; SRC="$(pwd)"
git clone https://huggingface.co/spaces/$USER/$SPACE ~/$SPACE-space && \
cp "$SRC/app.py" "$SRC/requirements.txt" "$SRC/deploy.md" "$SRC/deploy.ja.md" ~/$SPACE-space/ && \
cp "$SRC/README.gradio.md" ~/$SPACE-space/README.md && \
cd ~/$SPACE-space && git add app.py requirements.txt README.md deploy.md deploy.ja.md && \
git commit -m "Deploy EchoCache: semantic cache + cost observability" && git push
```
If `git clone` asks for credentials, use your account name as the username and your **access token** as the password.
The Space builds automatically (**Building** → **Running**, 1–2 minutes on a first push). Watch **Logs** on the Space page.
> Note the card swap: the published `README.md` carries `sdk: static`, so a Gradio deployment must use `README.gradio.md`.
> If the build complains about `sdk_version`, change that line in `README.md` to a version offered in your Space's Settings, then push again. Nothing else needs to change.
---
## 3. Set `HF_TOKEN` (optional — everything works without it)
Only needed if you want the optional embedding re-rank.
1. Space page → **Settings** → **Variables and secrets**
2. **New secret** → Name: `HF_TOKEN` / Value: your token → **Save**
(`HUGGINGFACE_HUB_TOKEN` and `HUGGINGFACEHUB_API_TOKEN` are read as well. If one of them is exported in your shell during a local run, you will start in `local+embedding` mode and make outbound calls — the header and the `Health` tab always say which mode is active.)
3. **Restart this Space**
| | without `HF_TOKEN` | with `HF_TOKEN` |
|---|---|---|
| similarity | fully local (default, recommended) | local + embedding, weighted average |
| core features | **all working** | all working |
| on failure | — | 402/429/timeout falls back to local automatically |
A free Hugging Face account has roughly **$0.10/month** of inference credit. **Running without the token is the normal configuration.**
Useful optional variables, same screen:
```
PRICE_IN_PER_1K = 0.003 # your input price; leave unset and savings stay $0
PRICE_OUT_PER_1K = 0.015 # your output price
MAX_ENTRIES = 5000 # per tenant
MAX_TOTAL_ENTRIES = 20000 # across all tenants — the real memory guard
MAX_TENANTS = 256 # partition cap; exceeding it drops the LRU tenant's whole partition
DEFAULT_THRESHOLD = 0.92 # conservative by default; move it after running a Sweep
```
---
## 4. Verify it works (3 minutes in the UI)
1. **Open the Space** (`https://<user>-echocache.hf.space`). The header should read `Mode: local-only`.
2. **Store one entry** — `Store` tab, `tenant_id` = `acme`, `route` = `support`,
prompt `How do I reset my password?`, response `Open Settings > Security > Reset password.` → **Store** returns `"stored": true` and an `entry_id`.
(Or press *Seed 3 demo entries*.)
3. **Look it up**
* `Lookup` tab, `tenant_id` = `acme`, prompt `Hi team, how do I reset my password? Thanks in advance!` → `"hit": true, "stage": "exact"`. The greeting and sign-off were normalized away onto the same key.
* `How can I reset my password?` at threshold `0.80` → `"stage": "cosine"` (measured similarity ≈ **0.85**; at the default `0.92` it misses and returns `best_similarity` 0.85 — **read that number before choosing a threshold**).
* **Wrong-hit protection:** `Why do I need to reset my password?` at threshold `0.70` → `"reason": "guard_rejected"` with `question_type_mismatch:q={why},c={how}`. High similarity is not enough.
* **Tenant isolation:** switch `tenant_id` to `globex` and repeat → `"reason": "empty_index"`. Another tenant's entries are unreachable.
4. **Dashboard** → **Refresh**: one-line summary (requests / hits / misses / hit rate / saved tokens / estimated saving), hit-rate time series, cumulative savings per tenant, p90 latency, most-reused entries, index health. **Export events CSV** takes the log with you (prompt text is not recorded by default).
5. **Sweep** → **Download a sample log**, upload that same file, enter your prices, **Run sweep**. You get "at this threshold you reuse X% and save $Y", with MismatchGuard **on** and **off** side by side — the conservative row is the one you can defend in a review.
6. **API Docs** tab shows the real paths for your URL and your Gradio version. The **Use via API** link in the page footer is generated by Gradio itself and is always authoritative.
### From the command line
```bash
SPACE="https://<user>-echocache.hf.space"
EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/store" -H "Content-Type: application/json" \
-d '{"data": ["acme","How do I reset my password?","Open Settings > Security > Reset password.","support",86400]}' \
| python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])")
curl -s -N "$SPACE/gradio_api/call/store/$EVENT_ID"
EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/lookup" -H "Content-Type: application/json" \
-d '{"data": ["acme","Hi team, how do I reset my password? Thanks!",0.92,"support"]}' \
| python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])")
curl -s -N "$SPACE/gradio_api/call/lookup/$EVENT_ID"
```
```bash
pip install gradio_client
python3 - <<'PY'
from gradio_client import Client
c = Client("<user>/EchoCache")
print(c.predict("acme", "how do i reset my password", 0.92, "support", api_name="/lookup"))
print(c.predict(api_name="/health"))
PY
```
---
## 5. How to run it
### 5-1. Export before the Space sleeps, import after it wakes
The free tier has **no persistent disk** and **sleeps after 48h of inactivity** (the timer cannot be changed). The index lives in RAM only, so carrying it over is an explicit step:
1. `Backup` tab → tick *include response bodies* → **Export index**
2. Save `index.json` somewhere outside the Space
3. After a sleep or restart: `Backup` tab → upload `index.json` → **Import index**
4. Check `imported / skipped_no_response / expired / rejected_by_safety` in the result
* Vectors are **recomputed** on import, so a changed `VECTOR_DIM` is not a problem.
* The safety screen **runs again** on import (leave `skip_safety` off).
* An export without response bodies is for analysis only — those rows are skipped on import and counted under `skipped_no_response`.
Daily backup from your own machine:
```bash
python3 - <<'PY'
from gradio_client import Client
import datetime
c = Client("<user>/EchoCache")
path, receipt = c.predict(True, "", api_name="/export_index")
dst = "echocache-%s.json" % datetime.date.today()
open(dst, "w").write(open(path).read())
print(dst, receipt)
PY
```
### 5-2. Choose the threshold from a Sweep, not from intuition
* The default `0.92` is **conservative**: it mostly reuses exact matches and surface variations (greetings, honorifics, punctuation, casing).
* Genuine paraphrases with different wording land around **0.80–0.90** with character n-grams.
* Export a day of real traffic as CSV (`tenant_id, prompt, response, route`), run **Sweep**, and read reuse rate, estimated saving, borderline count and guard blocks before deciding.
* Before lowering the threshold, open **Audit** and read why the guard rejected candidates. Many rejections mean "these questions genuinely differ", not "the threshold is too high".
* On the published benchmark, threshold **0.80 with the guard on** scored 99.2% accuracy with 100% reuse recall and 1 wrong reuse in 90 — but that is one synthetic set, not your traffic. Measure yours.
### 5-3. What to watch
| Frequency | Where | What you are deciding |
|---|---|---|
| daily | Dashboard hit rate / saved cost | is reuse growing as expected |
| daily | Audit → guard rejections | rejection patterns; early warning of wrong hits |
| weekly | Dashboard index health (`evicted` / `expired`) | revisit TTL or the entry caps |
| weekly | Sweep on fresh logs | revisit the threshold |
| on change | Health → `mode` | did it silently switch away from `local-only` |
---
## 6. Troubleshooting
### Memory pressure, or the Space restarts
Index memory is roughly `entries × VECTOR_DIM × 4 bytes`, and **`MAX_ENTRIES` is per tenant**. With the defaults, a full tenant is ≈ 80 MB and the whole index is capped at ≈ 330 MB by `MAX_TOTAL_ENTRIES=20000`.
In order of effect (Settings → Variables and secrets):
```
MAX_TOTAL_ENTRIES = 8000 # the global ceiling — start here
MAX_ENTRIES = 2000 # per tenant
MAX_TEXT_CHARS = 4000 # per stored prompt / response
VECTOR_DIM = 2048 # halves memory, costs little accuracy
EVENT_BUFFER = 2000 # event ring buffer
```
`Health` shows the live `approx_memory_mb` and `vector_matrix_mb_per_full_tenant`.
### A tenant's cache went empty
Reaching `MAX_TENANTS` (default 256) **evicts the least-recently-used tenant's entire partition** on the next new tenant's first write. Deletion only — nothing is ever moved or read across tenants. Size `MAX_TENANTS` above your real tenant count. When the global entry budget is reached instead, entries are evicted one at a time from the largest partition.
### Nothing hits
Work through this in order before touching the threshold:
1. Read `reason` in the Lookup result:
* `empty_index` — nothing stored for that tenant (a mistyped `tenant_id` is the usual cause)
* `no_candidates` — no fingerprint-similar entry; the wording differs a lot
* `below_threshold` — `best_similarity` is in the payload; decide from that number
* `guard_rejected` — the meaning was judged different; go to step 2
2. Open **Audit** and read the reasons (`numeric_mismatch`, `negation_mismatch`, `proper_noun_mismatch`, `temporal_mismatch`, `question_type_mismatch`). If the reason is sound, that miss is **correct** and lowering the threshold will not fix it.
3. Compare the stored and queried prompts: `cached_prompt_preview` comes back in the response.
4. Still missing legitimate reuse? Step down through Sweep (0.86 → 0.82) and watch `borderline_hits` and `guard_blocks` grow.
5. Only if a specific check genuinely does not fit your domain, disable it individually: `GUARD_DISABLE=proper_noun`. Disabling all of them is not recommended.
### A wrong hit happened (deal with this first)
1. **Raise the threshold** (e.g. `0.92` → `0.95`). Immediate effect.
2. Find the `entry_id` in **Audit**, or in Dashboard → *most reused entries*.
3. Use **Invalidate entry** at the bottom of the Audit tab (`tenant_id` + the first 16 characters of the id). A route name invalidates that route; `*` clears the tenant.
4. If the same shape recurs, add those prompts to a CSV, re-run Sweep, and adopt a threshold that rejects them.
5. If the difference is not one the five checks can see, stop caching that category: give it its own `route` and invalidate that route on a schedule.
### The embedding API returns 402 or 429
* `402 Payment Required` — the free inference credit (~$0.10/month) is exhausted.
* `429 Too Many Requests` — rate limited.
**Neither stops the cache.** The exception is swallowed and matching continues on local vectors alone. After repeated failures a circuit breaker opens for `EMBED_COOLDOWN_SEC` (default 600s) so no further credit is spent. `Health` shows `embedding.last_error` and `cooldown_remaining_sec`. To stop permanently, delete `HF_TOKEN` from Secrets and restart — the header returns to `local-only`. If it does not, check for `HUGGINGFACE_HUB_TOKEN` / `HUGGINGFACEHUB_API_TOKEN`.
### The Space sleeps after 48h and wakes up empty
That is the free tier working as designed (no persistent disk, fixed sleep timer). Use the export/import routine in 5-1. If you need it always on, move to paid hardware or run `python app.py` on your own host — the same file works unchanged.
### `/store` returns `stored: false`
The safety screen refused the content. Read `reason`:
| reason | meaning |
|---|---|
| `credit_card_luhn` | a digit run passed the Luhn checksum; order numbers are not flagged |
| `credential_prefix` | `sk-`, `ghp_`, `AKIA`, `xoxb-`, `AIza`, `hf_`, … |
| `jwt_structure` | a three-part JWT |
| `high_entropy_secret` | a long high-entropy run that is not identifier-shaped |
| `contact_pii_excess` | too many e-mail addresses or phone numbers |
| `private_key_block` | a PEM private-key block |
**This is correct behaviour.** Before relaxing anything, ask whether that content belongs in a cache at all.
### The API path 404s
Gradio changes it between major versions (5 and 6 use `/gradio_api/call/<name>`, 4 uses `/call/<name>`). Do not hardcode it — read it from the **API Docs** tab or the **Use via API** footer link. `gradio_client` resolves it for you.
---
## Appendix: running locally
```bash
pip install -r requirements.txt
python3 app.py # http://localhost:7860
```
Identical behaviour to the Space. Environment variables work the same way:
```bash
PRICE_IN_PER_1K=0.003 PRICE_OUT_PER_1K=0.015 MAX_ENTRIES=1000 python3 app.py
```