File size: 13,979 Bytes
46985e0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 | # Deploying EchoCache on a free Hugging Face Space
> 日本語版: [`deploy.ja.md`](deploy.ja.md)
About 10 minutes, copy-paste throughout. You need a browser, `git` and `python3`.
No paid plan, no GPU, no model download.
---
## 1. Create the Space
1. Open <https://huggingface.co/new-space>
2. Fill in:
| Field | Value |
|---|---|
| **Owner** | your account |
| **Space name** | `EchoCache` (referred to below as `<user>/EchoCache`) |
| **License** | `apache-2.0` |
| **SDK** | **Gradio** |
| **Space hardware** | **CPU basic · 2 vCPU · 16 GB · FREE** |
| **Visibility** | Public or Private — both work |
3. Press **Create Space**. If a Gradio template picker appears, choose "Blank"; every file gets replaced anyway.
---
## 2. Push the files
Log in once with a **write** token from <https://huggingface.co/settings/tokens>:
```bash
hf auth login # older installs: huggingface-cli login
```
Then run this as-is, replacing `<user>` with your account name:
```bash
USER=<user>; SPACE=EchoCache; SRC="$(pwd)"
git clone https://huggingface.co/spaces/$USER/$SPACE ~/$SPACE-space && \
cp "$SRC/app.py" "$SRC/requirements.txt" "$SRC/deploy.md" "$SRC/deploy.ja.md" ~/$SPACE-space/ && \
cp "$SRC/README.gradio.md" ~/$SPACE-space/README.md && \
cd ~/$SPACE-space && git add app.py requirements.txt README.md deploy.md deploy.ja.md && \
git commit -m "Deploy EchoCache: semantic cache + cost observability" && git push
```
If `git clone` asks for credentials, use your account name as the username and your **access token** as the password.
The Space builds automatically (**Building** → **Running**, 1–2 minutes on a first push). Watch **Logs** on the Space page.
> Note the card swap: the published `README.md` carries `sdk: static`, so a Gradio deployment must use `README.gradio.md`.
> If the build complains about `sdk_version`, change that line in `README.md` to a version offered in your Space's Settings, then push again. Nothing else needs to change.
---
## 3. Set `HF_TOKEN` (optional — everything works without it)
Only needed if you want the optional embedding re-rank.
1. Space page → **Settings** → **Variables and secrets**
2. **New secret** → Name: `HF_TOKEN` / Value: your token → **Save**
(`HUGGINGFACE_HUB_TOKEN` and `HUGGINGFACEHUB_API_TOKEN` are read as well. If one of them is exported in your shell during a local run, you will start in `local+embedding` mode and make outbound calls — the header and the `Health` tab always say which mode is active.)
3. **Restart this Space**
| | without `HF_TOKEN` | with `HF_TOKEN` |
|---|---|---|
| similarity | fully local (default, recommended) | local + embedding, weighted average |
| core features | **all working** | all working |
| on failure | — | 402/429/timeout falls back to local automatically |
A free Hugging Face account has roughly **$0.10/month** of inference credit. **Running without the token is the normal configuration.**
Useful optional variables, same screen:
```
PRICE_IN_PER_1K = 0.003 # your input price; leave unset and savings stay $0
PRICE_OUT_PER_1K = 0.015 # your output price
MAX_ENTRIES = 5000 # per tenant
MAX_TOTAL_ENTRIES = 20000 # across all tenants — the real memory guard
MAX_TENANTS = 256 # partition cap; exceeding it drops the LRU tenant's whole partition
DEFAULT_THRESHOLD = 0.92 # conservative by default; move it after running a Sweep
```
---
## 4. Verify it works (3 minutes in the UI)
1. **Open the Space** (`https://<user>-echocache.hf.space`). The header should read `Mode: local-only`.
2. **Store one entry** — `Store` tab, `tenant_id` = `acme`, `route` = `support`,
prompt `How do I reset my password?`, response `Open Settings > Security > Reset password.` → **Store** returns `"stored": true` and an `entry_id`.
(Or press *Seed 3 demo entries*.)
3. **Look it up**
* `Lookup` tab, `tenant_id` = `acme`, prompt `Hi team, how do I reset my password? Thanks in advance!` → `"hit": true, "stage": "exact"`. The greeting and sign-off were normalized away onto the same key.
* `How can I reset my password?` at threshold `0.80` → `"stage": "cosine"` (measured similarity ≈ **0.85**; at the default `0.92` it misses and returns `best_similarity` 0.85 — **read that number before choosing a threshold**).
* **Wrong-hit protection:** `Why do I need to reset my password?` at threshold `0.70` → `"reason": "guard_rejected"` with `question_type_mismatch:q={why},c={how}`. High similarity is not enough.
* **Tenant isolation:** switch `tenant_id` to `globex` and repeat → `"reason": "empty_index"`. Another tenant's entries are unreachable.
4. **Dashboard** → **Refresh**: one-line summary (requests / hits / misses / hit rate / saved tokens / estimated saving), hit-rate time series, cumulative savings per tenant, p90 latency, most-reused entries, index health. **Export events CSV** takes the log with you (prompt text is not recorded by default).
5. **Sweep** → **Download a sample log**, upload that same file, enter your prices, **Run sweep**. You get "at this threshold you reuse X% and save $Y", with MismatchGuard **on** and **off** side by side — the conservative row is the one you can defend in a review.
6. **API Docs** tab shows the real paths for your URL and your Gradio version. The **Use via API** link in the page footer is generated by Gradio itself and is always authoritative.
### From the command line
```bash
SPACE="https://<user>-echocache.hf.space"
EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/store" -H "Content-Type: application/json" \
-d '{"data": ["acme","How do I reset my password?","Open Settings > Security > Reset password.","support",86400]}' \
| python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])")
curl -s -N "$SPACE/gradio_api/call/store/$EVENT_ID"
EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/lookup" -H "Content-Type: application/json" \
-d '{"data": ["acme","Hi team, how do I reset my password? Thanks!",0.92,"support"]}' \
| python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])")
curl -s -N "$SPACE/gradio_api/call/lookup/$EVENT_ID"
```
```bash
pip install gradio_client
python3 - <<'PY'
from gradio_client import Client
c = Client("<user>/EchoCache")
print(c.predict("acme", "how do i reset my password", 0.92, "support", api_name="/lookup"))
print(c.predict(api_name="/health"))
PY
```
---
## 5. How to run it
### 5-1. Export before the Space sleeps, import after it wakes
The free tier has **no persistent disk** and **sleeps after 48h of inactivity** (the timer cannot be changed). The index lives in RAM only, so carrying it over is an explicit step:
1. `Backup` tab → tick *include response bodies* → **Export index**
2. Save `index.json` somewhere outside the Space
3. After a sleep or restart: `Backup` tab → upload `index.json` → **Import index**
4. Check `imported / skipped_no_response / expired / rejected_by_safety` in the result
* Vectors are **recomputed** on import, so a changed `VECTOR_DIM` is not a problem.
* The safety screen **runs again** on import (leave `skip_safety` off).
* An export without response bodies is for analysis only — those rows are skipped on import and counted under `skipped_no_response`.
Daily backup from your own machine:
```bash
python3 - <<'PY'
from gradio_client import Client
import datetime
c = Client("<user>/EchoCache")
path, receipt = c.predict(True, "", api_name="/export_index")
dst = "echocache-%s.json" % datetime.date.today()
open(dst, "w").write(open(path).read())
print(dst, receipt)
PY
```
### 5-2. Choose the threshold from a Sweep, not from intuition
* The default `0.92` is **conservative**: it mostly reuses exact matches and surface variations (greetings, honorifics, punctuation, casing).
* Genuine paraphrases with different wording land around **0.80–0.90** with character n-grams.
* Export a day of real traffic as CSV (`tenant_id, prompt, response, route`), run **Sweep**, and read reuse rate, estimated saving, borderline count and guard blocks before deciding.
* Before lowering the threshold, open **Audit** and read why the guard rejected candidates. Many rejections mean "these questions genuinely differ", not "the threshold is too high".
* On the published benchmark, threshold **0.80 with the guard on** scored 99.2% accuracy with 100% reuse recall and 1 wrong reuse in 90 — but that is one synthetic set, not your traffic. Measure yours.
### 5-3. What to watch
| Frequency | Where | What you are deciding |
|---|---|---|
| daily | Dashboard hit rate / saved cost | is reuse growing as expected |
| daily | Audit → guard rejections | rejection patterns; early warning of wrong hits |
| weekly | Dashboard index health (`evicted` / `expired`) | revisit TTL or the entry caps |
| weekly | Sweep on fresh logs | revisit the threshold |
| on change | Health → `mode` | did it silently switch away from `local-only` |
---
## 6. Troubleshooting
### Memory pressure, or the Space restarts
Index memory is roughly `entries × VECTOR_DIM × 4 bytes`, and **`MAX_ENTRIES` is per tenant**. With the defaults, a full tenant is ≈ 80 MB and the whole index is capped at ≈ 330 MB by `MAX_TOTAL_ENTRIES=20000`.
In order of effect (Settings → Variables and secrets):
```
MAX_TOTAL_ENTRIES = 8000 # the global ceiling — start here
MAX_ENTRIES = 2000 # per tenant
MAX_TEXT_CHARS = 4000 # per stored prompt / response
VECTOR_DIM = 2048 # halves memory, costs little accuracy
EVENT_BUFFER = 2000 # event ring buffer
```
`Health` shows the live `approx_memory_mb` and `vector_matrix_mb_per_full_tenant`.
### A tenant's cache went empty
Reaching `MAX_TENANTS` (default 256) **evicts the least-recently-used tenant's entire partition** on the next new tenant's first write. Deletion only — nothing is ever moved or read across tenants. Size `MAX_TENANTS` above your real tenant count. When the global entry budget is reached instead, entries are evicted one at a time from the largest partition.
### Nothing hits
Work through this in order before touching the threshold:
1. Read `reason` in the Lookup result:
* `empty_index` — nothing stored for that tenant (a mistyped `tenant_id` is the usual cause)
* `no_candidates` — no fingerprint-similar entry; the wording differs a lot
* `below_threshold` — `best_similarity` is in the payload; decide from that number
* `guard_rejected` — the meaning was judged different; go to step 2
2. Open **Audit** and read the reasons (`numeric_mismatch`, `negation_mismatch`, `proper_noun_mismatch`, `temporal_mismatch`, `question_type_mismatch`). If the reason is sound, that miss is **correct** and lowering the threshold will not fix it.
3. Compare the stored and queried prompts: `cached_prompt_preview` comes back in the response.
4. Still missing legitimate reuse? Step down through Sweep (0.86 → 0.82) and watch `borderline_hits` and `guard_blocks` grow.
5. Only if a specific check genuinely does not fit your domain, disable it individually: `GUARD_DISABLE=proper_noun`. Disabling all of them is not recommended.
### A wrong hit happened (deal with this first)
1. **Raise the threshold** (e.g. `0.92` → `0.95`). Immediate effect.
2. Find the `entry_id` in **Audit**, or in Dashboard → *most reused entries*.
3. Use **Invalidate entry** at the bottom of the Audit tab (`tenant_id` + the first 16 characters of the id). A route name invalidates that route; `*` clears the tenant.
4. If the same shape recurs, add those prompts to a CSV, re-run Sweep, and adopt a threshold that rejects them.
5. If the difference is not one the five checks can see, stop caching that category: give it its own `route` and invalidate that route on a schedule.
### The embedding API returns 402 or 429
* `402 Payment Required` — the free inference credit (~$0.10/month) is exhausted.
* `429 Too Many Requests` — rate limited.
**Neither stops the cache.** The exception is swallowed and matching continues on local vectors alone. After repeated failures a circuit breaker opens for `EMBED_COOLDOWN_SEC` (default 600s) so no further credit is spent. `Health` shows `embedding.last_error` and `cooldown_remaining_sec`. To stop permanently, delete `HF_TOKEN` from Secrets and restart — the header returns to `local-only`. If it does not, check for `HUGGINGFACE_HUB_TOKEN` / `HUGGINGFACEHUB_API_TOKEN`.
### The Space sleeps after 48h and wakes up empty
That is the free tier working as designed (no persistent disk, fixed sleep timer). Use the export/import routine in 5-1. If you need it always on, move to paid hardware or run `python app.py` on your own host — the same file works unchanged.
### `/store` returns `stored: false`
The safety screen refused the content. Read `reason`:
| reason | meaning |
|---|---|
| `credit_card_luhn` | a digit run passed the Luhn checksum; order numbers are not flagged |
| `credential_prefix` | `sk-`, `ghp_`, `AKIA`, `xoxb-`, `AIza`, `hf_`, … |
| `jwt_structure` | a three-part JWT |
| `high_entropy_secret` | a long high-entropy run that is not identifier-shaped |
| `contact_pii_excess` | too many e-mail addresses or phone numbers |
| `private_key_block` | a PEM private-key block |
**This is correct behaviour.** Before relaxing anything, ask whether that content belongs in a cache at all.
### The API path 404s
Gradio changes it between major versions (5 and 6 use `/gradio_api/call/<name>`, 4 uses `/call/<name>`). Do not hardcode it — read it from the **API Docs** tab or the **Use via API** footer link. `gradio_client` resolves it for you.
---
## Appendix: running locally
```bash
pip install -r requirements.txt
python3 app.py # http://localhost:7860
```
Identical behaviour to the Space. Environment variables work the same way:
```bash
PRICE_IN_PER_1K=0.003 PRICE_OUT_PER_1K=0.015 MAX_ENTRIES=1000 python3 app.py
```
|