# Deploying EchoCache on a free Hugging Face Space > 日本語版: [`deploy.ja.md`](deploy.ja.md) About 10 minutes, copy-paste throughout. You need a browser, `git` and `python3`. No paid plan, no GPU, no model download. --- ## 1. Create the Space 1. Open 2. Fill in: | Field | Value | |---|---| | **Owner** | your account | | **Space name** | `EchoCache` (referred to below as `/EchoCache`) | | **License** | `apache-2.0` | | **SDK** | **Gradio** | | **Space hardware** | **CPU basic · 2 vCPU · 16 GB · FREE** | | **Visibility** | Public or Private — both work | 3. Press **Create Space**. If a Gradio template picker appears, choose "Blank"; every file gets replaced anyway. --- ## 2. Push the files Log in once with a **write** token from : ```bash hf auth login # older installs: huggingface-cli login ``` Then run this as-is, replacing `` with your account name: ```bash USER=; SPACE=EchoCache; SRC="$(pwd)" git clone https://huggingface.co/spaces/$USER/$SPACE ~/$SPACE-space && \ cp "$SRC/app.py" "$SRC/requirements.txt" "$SRC/deploy.md" "$SRC/deploy.ja.md" ~/$SPACE-space/ && \ cp "$SRC/README.gradio.md" ~/$SPACE-space/README.md && \ cd ~/$SPACE-space && git add app.py requirements.txt README.md deploy.md deploy.ja.md && \ git commit -m "Deploy EchoCache: semantic cache + cost observability" && git push ``` If `git clone` asks for credentials, use your account name as the username and your **access token** as the password. The Space builds automatically (**Building** → **Running**, 1–2 minutes on a first push). Watch **Logs** on the Space page. > Note the card swap: the published `README.md` carries `sdk: static`, so a Gradio deployment must use `README.gradio.md`. > If the build complains about `sdk_version`, change that line in `README.md` to a version offered in your Space's Settings, then push again. Nothing else needs to change. --- ## 3. Set `HF_TOKEN` (optional — everything works without it) Only needed if you want the optional embedding re-rank. 1. Space page → **Settings** → **Variables and secrets** 2. **New secret** → Name: `HF_TOKEN` / Value: your token → **Save** (`HUGGINGFACE_HUB_TOKEN` and `HUGGINGFACEHUB_API_TOKEN` are read as well. If one of them is exported in your shell during a local run, you will start in `local+embedding` mode and make outbound calls — the header and the `Health` tab always say which mode is active.) 3. **Restart this Space** | | without `HF_TOKEN` | with `HF_TOKEN` | |---|---|---| | similarity | fully local (default, recommended) | local + embedding, weighted average | | core features | **all working** | all working | | on failure | — | 402/429/timeout falls back to local automatically | A free Hugging Face account has roughly **$0.10/month** of inference credit. **Running without the token is the normal configuration.** Useful optional variables, same screen: ``` PRICE_IN_PER_1K = 0.003 # your input price; leave unset and savings stay $0 PRICE_OUT_PER_1K = 0.015 # your output price MAX_ENTRIES = 5000 # per tenant MAX_TOTAL_ENTRIES = 20000 # across all tenants — the real memory guard MAX_TENANTS = 256 # partition cap; exceeding it drops the LRU tenant's whole partition DEFAULT_THRESHOLD = 0.92 # conservative by default; move it after running a Sweep ``` --- ## 4. Verify it works (3 minutes in the UI) 1. **Open the Space** (`https://-echocache.hf.space`). The header should read `Mode: local-only`. 2. **Store one entry** — `Store` tab, `tenant_id` = `acme`, `route` = `support`, prompt `How do I reset my password?`, response `Open Settings > Security > Reset password.` → **Store** returns `"stored": true` and an `entry_id`. (Or press *Seed 3 demo entries*.) 3. **Look it up** * `Lookup` tab, `tenant_id` = `acme`, prompt `Hi team, how do I reset my password? Thanks in advance!` → `"hit": true, "stage": "exact"`. The greeting and sign-off were normalized away onto the same key. * `How can I reset my password?` at threshold `0.80` → `"stage": "cosine"` (measured similarity ≈ **0.85**; at the default `0.92` it misses and returns `best_similarity` 0.85 — **read that number before choosing a threshold**). * **Wrong-hit protection:** `Why do I need to reset my password?` at threshold `0.70` → `"reason": "guard_rejected"` with `question_type_mismatch:q={why},c={how}`. High similarity is not enough. * **Tenant isolation:** switch `tenant_id` to `globex` and repeat → `"reason": "empty_index"`. Another tenant's entries are unreachable. 4. **Dashboard** → **Refresh**: one-line summary (requests / hits / misses / hit rate / saved tokens / estimated saving), hit-rate time series, cumulative savings per tenant, p90 latency, most-reused entries, index health. **Export events CSV** takes the log with you (prompt text is not recorded by default). 5. **Sweep** → **Download a sample log**, upload that same file, enter your prices, **Run sweep**. You get "at this threshold you reuse X% and save $Y", with MismatchGuard **on** and **off** side by side — the conservative row is the one you can defend in a review. 6. **API Docs** tab shows the real paths for your URL and your Gradio version. The **Use via API** link in the page footer is generated by Gradio itself and is always authoritative. ### From the command line ```bash SPACE="https://-echocache.hf.space" EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/store" -H "Content-Type: application/json" \ -d '{"data": ["acme","How do I reset my password?","Open Settings > Security > Reset password.","support",86400]}' \ | python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])") curl -s -N "$SPACE/gradio_api/call/store/$EVENT_ID" EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/lookup" -H "Content-Type: application/json" \ -d '{"data": ["acme","Hi team, how do I reset my password? Thanks!",0.92,"support"]}' \ | python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])") curl -s -N "$SPACE/gradio_api/call/lookup/$EVENT_ID" ``` ```bash pip install gradio_client python3 - <<'PY' from gradio_client import Client c = Client("/EchoCache") print(c.predict("acme", "how do i reset my password", 0.92, "support", api_name="/lookup")) print(c.predict(api_name="/health")) PY ``` --- ## 5. How to run it ### 5-1. Export before the Space sleeps, import after it wakes The free tier has **no persistent disk** and **sleeps after 48h of inactivity** (the timer cannot be changed). The index lives in RAM only, so carrying it over is an explicit step: 1. `Backup` tab → tick *include response bodies* → **Export index** 2. Save `index.json` somewhere outside the Space 3. After a sleep or restart: `Backup` tab → upload `index.json` → **Import index** 4. Check `imported / skipped_no_response / expired / rejected_by_safety` in the result * Vectors are **recomputed** on import, so a changed `VECTOR_DIM` is not a problem. * The safety screen **runs again** on import (leave `skip_safety` off). * An export without response bodies is for analysis only — those rows are skipped on import and counted under `skipped_no_response`. Daily backup from your own machine: ```bash python3 - <<'PY' from gradio_client import Client import datetime c = Client("/EchoCache") path, receipt = c.predict(True, "", api_name="/export_index") dst = "echocache-%s.json" % datetime.date.today() open(dst, "w").write(open(path).read()) print(dst, receipt) PY ``` ### 5-2. Choose the threshold from a Sweep, not from intuition * The default `0.92` is **conservative**: it mostly reuses exact matches and surface variations (greetings, honorifics, punctuation, casing). * Genuine paraphrases with different wording land around **0.80–0.90** with character n-grams. * Export a day of real traffic as CSV (`tenant_id, prompt, response, route`), run **Sweep**, and read reuse rate, estimated saving, borderline count and guard blocks before deciding. * Before lowering the threshold, open **Audit** and read why the guard rejected candidates. Many rejections mean "these questions genuinely differ", not "the threshold is too high". * On the published benchmark, threshold **0.80 with the guard on** scored 99.2% accuracy with 100% reuse recall and 1 wrong reuse in 90 — but that is one synthetic set, not your traffic. Measure yours. ### 5-3. What to watch | Frequency | Where | What you are deciding | |---|---|---| | daily | Dashboard hit rate / saved cost | is reuse growing as expected | | daily | Audit → guard rejections | rejection patterns; early warning of wrong hits | | weekly | Dashboard index health (`evicted` / `expired`) | revisit TTL or the entry caps | | weekly | Sweep on fresh logs | revisit the threshold | | on change | Health → `mode` | did it silently switch away from `local-only` | --- ## 6. Troubleshooting ### Memory pressure, or the Space restarts Index memory is roughly `entries × VECTOR_DIM × 4 bytes`, and **`MAX_ENTRIES` is per tenant**. With the defaults, a full tenant is ≈ 80 MB and the whole index is capped at ≈ 330 MB by `MAX_TOTAL_ENTRIES=20000`. In order of effect (Settings → Variables and secrets): ``` MAX_TOTAL_ENTRIES = 8000 # the global ceiling — start here MAX_ENTRIES = 2000 # per tenant MAX_TEXT_CHARS = 4000 # per stored prompt / response VECTOR_DIM = 2048 # halves memory, costs little accuracy EVENT_BUFFER = 2000 # event ring buffer ``` `Health` shows the live `approx_memory_mb` and `vector_matrix_mb_per_full_tenant`. ### A tenant's cache went empty Reaching `MAX_TENANTS` (default 256) **evicts the least-recently-used tenant's entire partition** on the next new tenant's first write. Deletion only — nothing is ever moved or read across tenants. Size `MAX_TENANTS` above your real tenant count. When the global entry budget is reached instead, entries are evicted one at a time from the largest partition. ### Nothing hits Work through this in order before touching the threshold: 1. Read `reason` in the Lookup result: * `empty_index` — nothing stored for that tenant (a mistyped `tenant_id` is the usual cause) * `no_candidates` — no fingerprint-similar entry; the wording differs a lot * `below_threshold` — `best_similarity` is in the payload; decide from that number * `guard_rejected` — the meaning was judged different; go to step 2 2. Open **Audit** and read the reasons (`numeric_mismatch`, `negation_mismatch`, `proper_noun_mismatch`, `temporal_mismatch`, `question_type_mismatch`). If the reason is sound, that miss is **correct** and lowering the threshold will not fix it. 3. Compare the stored and queried prompts: `cached_prompt_preview` comes back in the response. 4. Still missing legitimate reuse? Step down through Sweep (0.86 → 0.82) and watch `borderline_hits` and `guard_blocks` grow. 5. Only if a specific check genuinely does not fit your domain, disable it individually: `GUARD_DISABLE=proper_noun`. Disabling all of them is not recommended. ### A wrong hit happened (deal with this first) 1. **Raise the threshold** (e.g. `0.92` → `0.95`). Immediate effect. 2. Find the `entry_id` in **Audit**, or in Dashboard → *most reused entries*. 3. Use **Invalidate entry** at the bottom of the Audit tab (`tenant_id` + the first 16 characters of the id). A route name invalidates that route; `*` clears the tenant. 4. If the same shape recurs, add those prompts to a CSV, re-run Sweep, and adopt a threshold that rejects them. 5. If the difference is not one the five checks can see, stop caching that category: give it its own `route` and invalidate that route on a schedule. ### The embedding API returns 402 or 429 * `402 Payment Required` — the free inference credit (~$0.10/month) is exhausted. * `429 Too Many Requests` — rate limited. **Neither stops the cache.** The exception is swallowed and matching continues on local vectors alone. After repeated failures a circuit breaker opens for `EMBED_COOLDOWN_SEC` (default 600s) so no further credit is spent. `Health` shows `embedding.last_error` and `cooldown_remaining_sec`. To stop permanently, delete `HF_TOKEN` from Secrets and restart — the header returns to `local-only`. If it does not, check for `HUGGINGFACE_HUB_TOKEN` / `HUGGINGFACEHUB_API_TOKEN`. ### The Space sleeps after 48h and wakes up empty That is the free tier working as designed (no persistent disk, fixed sleep timer). Use the export/import routine in 5-1. If you need it always on, move to paid hardware or run `python app.py` on your own host — the same file works unchanged. ### `/store` returns `stored: false` The safety screen refused the content. Read `reason`: | reason | meaning | |---|---| | `credit_card_luhn` | a digit run passed the Luhn checksum; order numbers are not flagged | | `credential_prefix` | `sk-`, `ghp_`, `AKIA`, `xoxb-`, `AIza`, `hf_`, … | | `jwt_structure` | a three-part JWT | | `high_entropy_secret` | a long high-entropy run that is not identifier-shaped | | `contact_pii_excess` | too many e-mail addresses or phone numbers | | `private_key_block` | a PEM private-key block | **This is correct behaviour.** Before relaxing anything, ask whether that content belongs in a cache at all. ### The API path 404s Gradio changes it between major versions (5 and 6 use `/gradio_api/call/`, 4 uses `/call/`). Do not hardcode it — read it from the **API Docs** tab or the **Use via API** footer link. `gradio_client` resolves it for you. --- ## Appendix: running locally ```bash pip install -r requirements.txt python3 app.py # http://localhost:7860 ``` Identical behaviour to the Space. Environment variables work the same way: ```bash PRICE_IN_PER_1K=0.003 PRICE_OUT_PER_1K=0.015 MAX_ENTRIES=1000 python3 app.py ```