EchoCache / deploy.md
NagaYu's picture
EchoCache v1.0.0: docs, benchmark results and the full application source
46985e0 verified
|
Raw History Blame Contribute Delete
14 kB

Deploying EchoCache on a free Hugging Face Space

日本語版: deploy.ja.md

About 10 minutes, copy-paste throughout. You need a browser, git and python3. No paid plan, no GPU, no model download.


1. Create the Space

  1. Open https://huggingface.co/new-space

  2. Fill in:

    Field Value
    Owner your account
    Space name EchoCache (referred to below as <user>/EchoCache)
    License apache-2.0
    SDK Gradio
    Space hardware CPU basic · 2 vCPU · 16 GB · FREE
    Visibility Public or Private — both work
  3. Press Create Space. If a Gradio template picker appears, choose "Blank"; every file gets replaced anyway.


2. Push the files

Log in once with a write token from https://huggingface.co/settings/tokens:

hf auth login          # older installs: huggingface-cli login

Then run this as-is, replacing <user> with your account name:

USER=<user>; SPACE=EchoCache; SRC="$(pwd)"
git clone https://huggingface.co/spaces/$USER/$SPACE ~/$SPACE-space && \
cp "$SRC/app.py" "$SRC/requirements.txt" "$SRC/deploy.md" "$SRC/deploy.ja.md" ~/$SPACE-space/ && \
cp "$SRC/README.gradio.md" ~/$SPACE-space/README.md && \
cd ~/$SPACE-space && git add app.py requirements.txt README.md deploy.md deploy.ja.md && \
git commit -m "Deploy EchoCache: semantic cache + cost observability" && git push

If git clone asks for credentials, use your account name as the username and your access token as the password.

The Space builds automatically (Building → Running, 1–2 minutes on a first push). Watch Logs on the Space page.

Note the card swap: the published README.md carries sdk: static, so a Gradio deployment must use README.gradio.md.

If the build complains about sdk_version, change that line in README.md to a version offered in your Space's Settings, then push again. Nothing else needs to change.


3. Set HF_TOKEN (optional — everything works without it)

Only needed if you want the optional embedding re-rank.

  1. Space page → Settings → Variables and secrets
  2. New secret → Name: HF_TOKEN / Value: your token → Save (HUGGINGFACE_HUB_TOKEN and HUGGINGFACEHUB_API_TOKEN are read as well. If one of them is exported in your shell during a local run, you will start in local+embedding mode and make outbound calls — the header and the Health tab always say which mode is active.)
  3. Restart this Space
without HF_TOKEN with HF_TOKEN
similarity fully local (default, recommended) local + embedding, weighted average
core features all working all working
on failure — 402/429/timeout falls back to local automatically

A free Hugging Face account has roughly $0.10/month of inference credit. Running without the token is the normal configuration.

Useful optional variables, same screen:

PRICE_IN_PER_1K   = 0.003     # your input price; leave unset and savings stay $0
PRICE_OUT_PER_1K  = 0.015     # your output price
MAX_ENTRIES       = 5000      # per tenant
MAX_TOTAL_ENTRIES = 20000     # across all tenants — the real memory guard
MAX_TENANTS       = 256       # partition cap; exceeding it drops the LRU tenant's whole partition
DEFAULT_THRESHOLD = 0.92      # conservative by default; move it after running a Sweep

4. Verify it works (3 minutes in the UI)

  1. Open the Space (https://<user>-echocache.hf.space). The header should read Mode: local-only.

  2. Store one entry — Store tab, tenant_id = acme, route = support, prompt How do I reset my password?, response Open Settings > Security > Reset password. → Store returns "stored": true and an entry_id. (Or press Seed 3 demo entries.)

  3. Look it up

    • Lookup tab, tenant_id = acme, prompt Hi team, how do I reset my password? Thanks in advance! → "hit": true, "stage": "exact". The greeting and sign-off were normalized away onto the same key.
    • How can I reset my password? at threshold 0.80 → "stage": "cosine" (measured similarity ≈ 0.85; at the default 0.92 it misses and returns best_similarity 0.85 — read that number before choosing a threshold).
    • Wrong-hit protection: Why do I need to reset my password? at threshold 0.70 → "reason": "guard_rejected" with question_type_mismatch:q={why},c={how}. High similarity is not enough.
    • Tenant isolation: switch tenant_id to globex and repeat → "reason": "empty_index". Another tenant's entries are unreachable.
  4. Dashboard → Refresh: one-line summary (requests / hits / misses / hit rate / saved tokens / estimated saving), hit-rate time series, cumulative savings per tenant, p90 latency, most-reused entries, index health. Export events CSV takes the log with you (prompt text is not recorded by default).

  5. Sweep → Download a sample log, upload that same file, enter your prices, Run sweep. You get "at this threshold you reuse X% and save $Y", with MismatchGuard on and off side by side — the conservative row is the one you can defend in a review.

  6. API Docs tab shows the real paths for your URL and your Gradio version. The Use via API link in the page footer is generated by Gradio itself and is always authoritative.

From the command line

SPACE="https://<user>-echocache.hf.space"

EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/store" -H "Content-Type: application/json" \
  -d '{"data": ["acme","How do I reset my password?","Open Settings > Security > Reset password.","support",86400]}' \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])")
curl -s -N "$SPACE/gradio_api/call/store/$EVENT_ID"

EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/lookup" -H "Content-Type: application/json" \
  -d '{"data": ["acme","Hi team, how do I reset my password? Thanks!",0.92,"support"]}' \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])")
curl -s -N "$SPACE/gradio_api/call/lookup/$EVENT_ID"
pip install gradio_client
python3 - <<'PY'
from gradio_client import Client
c = Client("<user>/EchoCache")
print(c.predict("acme", "how do i reset my password", 0.92, "support", api_name="/lookup"))
print(c.predict(api_name="/health"))
PY

5. How to run it

5-1. Export before the Space sleeps, import after it wakes

The free tier has no persistent disk and sleeps after 48h of inactivity (the timer cannot be changed). The index lives in RAM only, so carrying it over is an explicit step:

  1. Backup tab → tick include response bodies → Export index
  2. Save index.json somewhere outside the Space
  3. After a sleep or restart: Backup tab → upload index.json → Import index
  4. Check imported / skipped_no_response / expired / rejected_by_safety in the result
  • Vectors are recomputed on import, so a changed VECTOR_DIM is not a problem.
  • The safety screen runs again on import (leave skip_safety off).
  • An export without response bodies is for analysis only — those rows are skipped on import and counted under skipped_no_response.

Daily backup from your own machine:

python3 - <<'PY'
from gradio_client import Client
import datetime
c = Client("<user>/EchoCache")
path, receipt = c.predict(True, "", api_name="/export_index")
dst = "echocache-%s.json" % datetime.date.today()
open(dst, "w").write(open(path).read())
print(dst, receipt)
PY

5-2. Choose the threshold from a Sweep, not from intuition

  • The default 0.92 is conservative: it mostly reuses exact matches and surface variations (greetings, honorifics, punctuation, casing).
  • Genuine paraphrases with different wording land around 0.80–0.90 with character n-grams.
  • Export a day of real traffic as CSV (tenant_id, prompt, response, route), run Sweep, and read reuse rate, estimated saving, borderline count and guard blocks before deciding.
  • Before lowering the threshold, open Audit and read why the guard rejected candidates. Many rejections mean "these questions genuinely differ", not "the threshold is too high".
  • On the published benchmark, threshold 0.80 with the guard on scored 99.2% accuracy with 100% reuse recall and 1 wrong reuse in 90 — but that is one synthetic set, not your traffic. Measure yours.

5-3. What to watch

Frequency Where What you are deciding
daily Dashboard hit rate / saved cost is reuse growing as expected
daily Audit → guard rejections rejection patterns; early warning of wrong hits
weekly Dashboard index health (evicted / expired) revisit TTL or the entry caps
weekly Sweep on fresh logs revisit the threshold
on change Health → mode did it silently switch away from local-only

6. Troubleshooting

Memory pressure, or the Space restarts

Index memory is roughly entries × VECTOR_DIM × 4 bytes, and MAX_ENTRIES is per tenant. With the defaults, a full tenant is ≈ 80 MB and the whole index is capped at ≈ 330 MB by MAX_TOTAL_ENTRIES=20000.

In order of effect (Settings → Variables and secrets):

MAX_TOTAL_ENTRIES = 8000      # the global ceiling — start here
MAX_ENTRIES       = 2000      # per tenant
MAX_TEXT_CHARS    = 4000      # per stored prompt / response
VECTOR_DIM        = 2048      # halves memory, costs little accuracy
EVENT_BUFFER      = 2000      # event ring buffer

Health shows the live approx_memory_mb and vector_matrix_mb_per_full_tenant.

A tenant's cache went empty

Reaching MAX_TENANTS (default 256) evicts the least-recently-used tenant's entire partition on the next new tenant's first write. Deletion only — nothing is ever moved or read across tenants. Size MAX_TENANTS above your real tenant count. When the global entry budget is reached instead, entries are evicted one at a time from the largest partition.

Nothing hits

Work through this in order before touching the threshold:

  1. Read reason in the Lookup result:
    • empty_index — nothing stored for that tenant (a mistyped tenant_id is the usual cause)
    • no_candidates — no fingerprint-similar entry; the wording differs a lot
    • below_threshold — best_similarity is in the payload; decide from that number
    • guard_rejected — the meaning was judged different; go to step 2
  2. Open Audit and read the reasons (numeric_mismatch, negation_mismatch, proper_noun_mismatch, temporal_mismatch, question_type_mismatch). If the reason is sound, that miss is correct and lowering the threshold will not fix it.
  3. Compare the stored and queried prompts: cached_prompt_preview comes back in the response.
  4. Still missing legitimate reuse? Step down through Sweep (0.86 → 0.82) and watch borderline_hits and guard_blocks grow.
  5. Only if a specific check genuinely does not fit your domain, disable it individually: GUARD_DISABLE=proper_noun. Disabling all of them is not recommended.

A wrong hit happened (deal with this first)

  1. Raise the threshold (e.g. 0.92 → 0.95). Immediate effect.
  2. Find the entry_id in Audit, or in Dashboard → most reused entries.
  3. Use Invalidate entry at the bottom of the Audit tab (tenant_id + the first 16 characters of the id). A route name invalidates that route; * clears the tenant.
  4. If the same shape recurs, add those prompts to a CSV, re-run Sweep, and adopt a threshold that rejects them.
  5. If the difference is not one the five checks can see, stop caching that category: give it its own route and invalidate that route on a schedule.

The embedding API returns 402 or 429

  • 402 Payment Required — the free inference credit (~$0.10/month) is exhausted.
  • 429 Too Many Requests — rate limited.

Neither stops the cache. The exception is swallowed and matching continues on local vectors alone. After repeated failures a circuit breaker opens for EMBED_COOLDOWN_SEC (default 600s) so no further credit is spent. Health shows embedding.last_error and cooldown_remaining_sec. To stop permanently, delete HF_TOKEN from Secrets and restart — the header returns to local-only. If it does not, check for HUGGINGFACE_HUB_TOKEN / HUGGINGFACEHUB_API_TOKEN.

The Space sleeps after 48h and wakes up empty

That is the free tier working as designed (no persistent disk, fixed sleep timer). Use the export/import routine in 5-1. If you need it always on, move to paid hardware or run python app.py on your own host — the same file works unchanged.

/store returns stored: false

The safety screen refused the content. Read reason:

reason meaning
credit_card_luhn a digit run passed the Luhn checksum; order numbers are not flagged
credential_prefix sk-, ghp_, AKIA, xoxb-, AIza, hf_, …
jwt_structure a three-part JWT
high_entropy_secret a long high-entropy run that is not identifier-shaped
contact_pii_excess too many e-mail addresses or phone numbers
private_key_block a PEM private-key block

This is correct behaviour. Before relaxing anything, ask whether that content belongs in a cache at all.

The API path 404s

Gradio changes it between major versions (5 and 6 use /gradio_api/call/<name>, 4 uses /call/<name>). Do not hardcode it — read it from the API Docs tab or the Use via API footer link. gradio_client resolves it for you.


Appendix: running locally

pip install -r requirements.txt
python3 app.py          # http://localhost:7860

Identical behaviour to the Space. Environment variables work the same way:

PRICE_IN_PER_1K=0.003 PRICE_OUT_PER_1K=0.015 MAX_ENTRIES=1000 python3 app.py