|
Download deploy.md from NagaYu/EchoCache: direct link, hf CLI and curl.
- Browser
- Download file 14 kB
-
https://huggingface.co/spaces/NagaYu/EchoCache/resolve/main/deploy.md
- Command line
-
hf download hf://spaces/NagaYu/EchoCache/deploy.md
-
curl -L -o deploy.md https://huggingface.co/spaces/NagaYu/EchoCache/resolve/main/deploy.md
14 kB
| # Deploying EchoCache on a free Hugging Face Space | |
| > 日本語版: [`deploy.ja.md`](deploy.ja.md) | |
| About 10 minutes, copy-paste throughout. You need a browser, `git` and `python3`. | |
| No paid plan, no GPU, no model download. | |
| --- | |
| ## 1. Create the Space | |
| 1. Open <https://huggingface.co/new-space> | |
| 2. Fill in: | |
| | Field | Value | | |
| |---|---| | |
| | **Owner** | your account | | |
| | **Space name** | `EchoCache` (referred to below as `<user>/EchoCache`) | | |
| | **License** | `apache-2.0` | | |
| | **SDK** | **Gradio** | | |
| | **Space hardware** | **CPU basic · 2 vCPU · 16 GB · FREE** | | |
| | **Visibility** | Public or Private — both work | | |
| 3. Press **Create Space**. If a Gradio template picker appears, choose "Blank"; every file gets replaced anyway. | |
| --- | |
| ## 2. Push the files | |
| Log in once with a **write** token from <https://huggingface.co/settings/tokens>: | |
| ```bash | |
| hf auth login # older installs: huggingface-cli login | |
| ``` | |
| Then run this as-is, replacing `<user>` with your account name: | |
| ```bash | |
| USER=<user>; SPACE=EchoCache; SRC="$(pwd)" | |
| git clone https://huggingface.co/spaces/$USER/$SPACE ~/$SPACE-space && \ | |
| cp "$SRC/app.py" "$SRC/requirements.txt" "$SRC/deploy.md" "$SRC/deploy.ja.md" ~/$SPACE-space/ && \ | |
| cp "$SRC/README.gradio.md" ~/$SPACE-space/README.md && \ | |
| cd ~/$SPACE-space && git add app.py requirements.txt README.md deploy.md deploy.ja.md && \ | |
| git commit -m "Deploy EchoCache: semantic cache + cost observability" && git push | |
| ``` | |
| If `git clone` asks for credentials, use your account name as the username and your **access token** as the password. | |
| The Space builds automatically (**Building** → **Running**, 1–2 minutes on a first push). Watch **Logs** on the Space page. | |
| > Note the card swap: the published `README.md` carries `sdk: static`, so a Gradio deployment must use `README.gradio.md`. | |
| > If the build complains about `sdk_version`, change that line in `README.md` to a version offered in your Space's Settings, then push again. Nothing else needs to change. | |
| --- | |
| ## 3. Set `HF_TOKEN` (optional — everything works without it) | |
| Only needed if you want the optional embedding re-rank. | |
| 1. Space page → **Settings** → **Variables and secrets** | |
| 2. **New secret** → Name: `HF_TOKEN` / Value: your token → **Save** | |
| (`HUGGINGFACE_HUB_TOKEN` and `HUGGINGFACEHUB_API_TOKEN` are read as well. If one of them is exported in your shell during a local run, you will start in `local+embedding` mode and make outbound calls — the header and the `Health` tab always say which mode is active.) | |
| 3. **Restart this Space** | |
| | | without `HF_TOKEN` | with `HF_TOKEN` | | |
| |---|---|---| | |
| | similarity | fully local (default, recommended) | local + embedding, weighted average | | |
| | core features | **all working** | all working | | |
| | on failure | — | 402/429/timeout falls back to local automatically | | |
| A free Hugging Face account has roughly **$0.10/month** of inference credit. **Running without the token is the normal configuration.** | |
| Useful optional variables, same screen: | |
| ``` | |
| PRICE_IN_PER_1K = 0.003 # your input price; leave unset and savings stay $0 | |
| PRICE_OUT_PER_1K = 0.015 # your output price | |
| MAX_ENTRIES = 5000 # per tenant | |
| MAX_TOTAL_ENTRIES = 20000 # across all tenants — the real memory guard | |
| MAX_TENANTS = 256 # partition cap; exceeding it drops the LRU tenant's whole partition | |
| DEFAULT_THRESHOLD = 0.92 # conservative by default; move it after running a Sweep | |
| ``` | |
| --- | |
| ## 4. Verify it works (3 minutes in the UI) | |
| 1. **Open the Space** (`https://<user>-echocache.hf.space`). The header should read `Mode: local-only`. | |
| 2. **Store one entry** — `Store` tab, `tenant_id` = `acme`, `route` = `support`, | |
| prompt `How do I reset my password?`, response `Open Settings > Security > Reset password.` → **Store** returns `"stored": true` and an `entry_id`. | |
| (Or press *Seed 3 demo entries*.) | |
| 3. **Look it up** | |
| * `Lookup` tab, `tenant_id` = `acme`, prompt `Hi team, how do I reset my password? Thanks in advance!` → `"hit": true, "stage": "exact"`. The greeting and sign-off were normalized away onto the same key. | |
| * `How can I reset my password?` at threshold `0.80` → `"stage": "cosine"` (measured similarity ≈ **0.85**; at the default `0.92` it misses and returns `best_similarity` 0.85 — **read that number before choosing a threshold**). | |
| * **Wrong-hit protection:** `Why do I need to reset my password?` at threshold `0.70` → `"reason": "guard_rejected"` with `question_type_mismatch:q={why},c={how}`. High similarity is not enough. | |
| * **Tenant isolation:** switch `tenant_id` to `globex` and repeat → `"reason": "empty_index"`. Another tenant's entries are unreachable. | |
| 4. **Dashboard** → **Refresh**: one-line summary (requests / hits / misses / hit rate / saved tokens / estimated saving), hit-rate time series, cumulative savings per tenant, p90 latency, most-reused entries, index health. **Export events CSV** takes the log with you (prompt text is not recorded by default). | |
| 5. **Sweep** → **Download a sample log**, upload that same file, enter your prices, **Run sweep**. You get "at this threshold you reuse X% and save $Y", with MismatchGuard **on** and **off** side by side — the conservative row is the one you can defend in a review. | |
| 6. **API Docs** tab shows the real paths for your URL and your Gradio version. The **Use via API** link in the page footer is generated by Gradio itself and is always authoritative. | |
| ### From the command line | |
| ```bash | |
| SPACE="https://<user>-echocache.hf.space" | |
| EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/store" -H "Content-Type: application/json" \ | |
| -d '{"data": ["acme","How do I reset my password?","Open Settings > Security > Reset password.","support",86400]}' \ | |
| | python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])") | |
| curl -s -N "$SPACE/gradio_api/call/store/$EVENT_ID" | |
| EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/lookup" -H "Content-Type: application/json" \ | |
| -d '{"data": ["acme","Hi team, how do I reset my password? Thanks!",0.92,"support"]}' \ | |
| | python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])") | |
| curl -s -N "$SPACE/gradio_api/call/lookup/$EVENT_ID" | |
| ``` | |
| ```bash | |
| pip install gradio_client | |
| python3 - <<'PY' | |
| from gradio_client import Client | |
| c = Client("<user>/EchoCache") | |
| print(c.predict("acme", "how do i reset my password", 0.92, "support", api_name="/lookup")) | |
| print(c.predict(api_name="/health")) | |
| PY | |
| ``` | |
| --- | |
| ## 5. How to run it | |
| ### 5-1. Export before the Space sleeps, import after it wakes | |
| The free tier has **no persistent disk** and **sleeps after 48h of inactivity** (the timer cannot be changed). The index lives in RAM only, so carrying it over is an explicit step: | |
| 1. `Backup` tab → tick *include response bodies* → **Export index** | |
| 2. Save `index.json` somewhere outside the Space | |
| 3. After a sleep or restart: `Backup` tab → upload `index.json` → **Import index** | |
| 4. Check `imported / skipped_no_response / expired / rejected_by_safety` in the result | |
| * Vectors are **recomputed** on import, so a changed `VECTOR_DIM` is not a problem. | |
| * The safety screen **runs again** on import (leave `skip_safety` off). | |
| * An export without response bodies is for analysis only — those rows are skipped on import and counted under `skipped_no_response`. | |
| Daily backup from your own machine: | |
| ```bash | |
| python3 - <<'PY' | |
| from gradio_client import Client | |
| import datetime | |
| c = Client("<user>/EchoCache") | |
| path, receipt = c.predict(True, "", api_name="/export_index") | |
| dst = "echocache-%s.json" % datetime.date.today() | |
| open(dst, "w").write(open(path).read()) | |
| print(dst, receipt) | |
| PY | |
| ``` | |
| ### 5-2. Choose the threshold from a Sweep, not from intuition | |
| * The default `0.92` is **conservative**: it mostly reuses exact matches and surface variations (greetings, honorifics, punctuation, casing). | |
| * Genuine paraphrases with different wording land around **0.80–0.90** with character n-grams. | |
| * Export a day of real traffic as CSV (`tenant_id, prompt, response, route`), run **Sweep**, and read reuse rate, estimated saving, borderline count and guard blocks before deciding. | |
| * Before lowering the threshold, open **Audit** and read why the guard rejected candidates. Many rejections mean "these questions genuinely differ", not "the threshold is too high". | |
| * On the published benchmark, threshold **0.80 with the guard on** scored 99.2% accuracy with 100% reuse recall and 1 wrong reuse in 90 — but that is one synthetic set, not your traffic. Measure yours. | |
| ### 5-3. What to watch | |
| | Frequency | Where | What you are deciding | | |
| |---|---|---| | |
| | daily | Dashboard hit rate / saved cost | is reuse growing as expected | | |
| | daily | Audit → guard rejections | rejection patterns; early warning of wrong hits | | |
| | weekly | Dashboard index health (`evicted` / `expired`) | revisit TTL or the entry caps | | |
| | weekly | Sweep on fresh logs | revisit the threshold | | |
| | on change | Health → `mode` | did it silently switch away from `local-only` | | |
| --- | |
| ## 6. Troubleshooting | |
| ### Memory pressure, or the Space restarts | |
| Index memory is roughly `entries × VECTOR_DIM × 4 bytes`, and **`MAX_ENTRIES` is per tenant**. With the defaults, a full tenant is ≈ 80 MB and the whole index is capped at ≈ 330 MB by `MAX_TOTAL_ENTRIES=20000`. | |
| In order of effect (Settings → Variables and secrets): | |
| ``` | |
| MAX_TOTAL_ENTRIES = 8000 # the global ceiling — start here | |
| MAX_ENTRIES = 2000 # per tenant | |
| MAX_TEXT_CHARS = 4000 # per stored prompt / response | |
| VECTOR_DIM = 2048 # halves memory, costs little accuracy | |
| EVENT_BUFFER = 2000 # event ring buffer | |
| ``` | |
| `Health` shows the live `approx_memory_mb` and `vector_matrix_mb_per_full_tenant`. | |
| ### A tenant's cache went empty | |
| Reaching `MAX_TENANTS` (default 256) **evicts the least-recently-used tenant's entire partition** on the next new tenant's first write. Deletion only — nothing is ever moved or read across tenants. Size `MAX_TENANTS` above your real tenant count. When the global entry budget is reached instead, entries are evicted one at a time from the largest partition. | |
| ### Nothing hits | |
| Work through this in order before touching the threshold: | |
| 1. Read `reason` in the Lookup result: | |
| * `empty_index` — nothing stored for that tenant (a mistyped `tenant_id` is the usual cause) | |
| * `no_candidates` — no fingerprint-similar entry; the wording differs a lot | |
| * `below_threshold` — `best_similarity` is in the payload; decide from that number | |
| * `guard_rejected` — the meaning was judged different; go to step 2 | |
| 2. Open **Audit** and read the reasons (`numeric_mismatch`, `negation_mismatch`, `proper_noun_mismatch`, `temporal_mismatch`, `question_type_mismatch`). If the reason is sound, that miss is **correct** and lowering the threshold will not fix it. | |
| 3. Compare the stored and queried prompts: `cached_prompt_preview` comes back in the response. | |
| 4. Still missing legitimate reuse? Step down through Sweep (0.86 → 0.82) and watch `borderline_hits` and `guard_blocks` grow. | |
| 5. Only if a specific check genuinely does not fit your domain, disable it individually: `GUARD_DISABLE=proper_noun`. Disabling all of them is not recommended. | |
| ### A wrong hit happened (deal with this first) | |
| 1. **Raise the threshold** (e.g. `0.92` → `0.95`). Immediate effect. | |
| 2. Find the `entry_id` in **Audit**, or in Dashboard → *most reused entries*. | |
| 3. Use **Invalidate entry** at the bottom of the Audit tab (`tenant_id` + the first 16 characters of the id). A route name invalidates that route; `*` clears the tenant. | |
| 4. If the same shape recurs, add those prompts to a CSV, re-run Sweep, and adopt a threshold that rejects them. | |
| 5. If the difference is not one the five checks can see, stop caching that category: give it its own `route` and invalidate that route on a schedule. | |
| ### The embedding API returns 402 or 429 | |
| * `402 Payment Required` — the free inference credit (~$0.10/month) is exhausted. | |
| * `429 Too Many Requests` — rate limited. | |
| **Neither stops the cache.** The exception is swallowed and matching continues on local vectors alone. After repeated failures a circuit breaker opens for `EMBED_COOLDOWN_SEC` (default 600s) so no further credit is spent. `Health` shows `embedding.last_error` and `cooldown_remaining_sec`. To stop permanently, delete `HF_TOKEN` from Secrets and restart — the header returns to `local-only`. If it does not, check for `HUGGINGFACE_HUB_TOKEN` / `HUGGINGFACEHUB_API_TOKEN`. | |
| ### The Space sleeps after 48h and wakes up empty | |
| That is the free tier working as designed (no persistent disk, fixed sleep timer). Use the export/import routine in 5-1. If you need it always on, move to paid hardware or run `python app.py` on your own host — the same file works unchanged. | |
| ### `/store` returns `stored: false` | |
| The safety screen refused the content. Read `reason`: | |
| | reason | meaning | | |
| |---|---| | |
| | `credit_card_luhn` | a digit run passed the Luhn checksum; order numbers are not flagged | | |
| | `credential_prefix` | `sk-`, `ghp_`, `AKIA`, `xoxb-`, `AIza`, `hf_`, … | | |
| | `jwt_structure` | a three-part JWT | | |
| | `high_entropy_secret` | a long high-entropy run that is not identifier-shaped | | |
| | `contact_pii_excess` | too many e-mail addresses or phone numbers | | |
| | `private_key_block` | a PEM private-key block | | |
| **This is correct behaviour.** Before relaxing anything, ask whether that content belongs in a cache at all. | |
| ### The API path 404s | |
| Gradio changes it between major versions (5 and 6 use `/gradio_api/call/<name>`, 4 uses `/call/<name>`). Do not hardcode it — read it from the **API Docs** tab or the **Use via API** footer link. `gradio_client` resolves it for you. | |
| --- | |
| ## Appendix: running locally | |
| ```bash | |
| pip install -r requirements.txt | |
| python3 app.py # http://localhost:7860 | |
| ``` | |
| Identical behaviour to the Space. Environment variables work the same way: | |
| ```bash | |
| PRICE_IN_PER_1K=0.003 PRICE_OUT_PER_1K=0.015 MAX_ENTRIES=1000 python3 app.py | |
| ``` | |