Download deploy.md from NagaYu/EchoCache: direct link, hf CLI and curl.
- Browser
- Download file 14 kB
-
https://huggingface.co/spaces/NagaYu/EchoCache/resolve/main/deploy.md
- Command line
-
hf download hf://spaces/NagaYu/EchoCache/deploy.md
-
curl -L -o deploy.md https://huggingface.co/spaces/NagaYu/EchoCache/resolve/main/deploy.md
Deploying EchoCache on a free Hugging Face Space
日本語版:
deploy.ja.md
About 10 minutes, copy-paste throughout. You need a browser, git and python3.
No paid plan, no GPU, no model download.
1. Create the Space
Fill in:
Field Value Owner your account Space name EchoCache(referred to below as<user>/EchoCache)License apache-2.0SDK Gradio Space hardware CPU basic · 2 vCPU · 16 GB · FREE Visibility Public or Private — both work Press Create Space. If a Gradio template picker appears, choose "Blank"; every file gets replaced anyway.
2. Push the files
Log in once with a write token from https://huggingface.co/settings/tokens:
hf auth login # older installs: huggingface-cli login
Then run this as-is, replacing <user> with your account name:
USER=<user>; SPACE=EchoCache; SRC="$(pwd)"
git clone https://huggingface.co/spaces/$USER/$SPACE ~/$SPACE-space && \
cp "$SRC/app.py" "$SRC/requirements.txt" "$SRC/deploy.md" "$SRC/deploy.ja.md" ~/$SPACE-space/ && \
cp "$SRC/README.gradio.md" ~/$SPACE-space/README.md && \
cd ~/$SPACE-space && git add app.py requirements.txt README.md deploy.md deploy.ja.md && \
git commit -m "Deploy EchoCache: semantic cache + cost observability" && git push
If git clone asks for credentials, use your account name as the username and your access token as the password.
The Space builds automatically (Building → Running, 1–2 minutes on a first push). Watch Logs on the Space page.
Note the card swap: the published
README.mdcarriessdk: static, so a Gradio deployment must useREADME.gradio.md.
If the build complains about
sdk_version, change that line inREADME.mdto a version offered in your Space's Settings, then push again. Nothing else needs to change.
3. Set HF_TOKEN (optional — everything works without it)
Only needed if you want the optional embedding re-rank.
- Space page → Settings → Variables and secrets
- New secret → Name:
HF_TOKEN/ Value: your token → Save (HUGGINGFACE_HUB_TOKENandHUGGINGFACEHUB_API_TOKENare read as well. If one of them is exported in your shell during a local run, you will start inlocal+embeddingmode and make outbound calls — the header and theHealthtab always say which mode is active.) - Restart this Space
without HF_TOKEN |
with HF_TOKEN |
|
|---|---|---|
| similarity | fully local (default, recommended) | local + embedding, weighted average |
| core features | all working | all working |
| on failure | — | 402/429/timeout falls back to local automatically |
A free Hugging Face account has roughly $0.10/month of inference credit. Running without the token is the normal configuration.
Useful optional variables, same screen:
PRICE_IN_PER_1K = 0.003 # your input price; leave unset and savings stay $0
PRICE_OUT_PER_1K = 0.015 # your output price
MAX_ENTRIES = 5000 # per tenant
MAX_TOTAL_ENTRIES = 20000 # across all tenants — the real memory guard
MAX_TENANTS = 256 # partition cap; exceeding it drops the LRU tenant's whole partition
DEFAULT_THRESHOLD = 0.92 # conservative by default; move it after running a Sweep
4. Verify it works (3 minutes in the UI)
Open the Space (
https://<user>-echocache.hf.space). The header should readMode: local-only.Store one entry —
Storetab,tenant_id=acme,route=support, promptHow do I reset my password?, responseOpen Settings > Security > Reset password.→ Store returns"stored": trueand anentry_id. (Or press Seed 3 demo entries.)Look it up
Lookuptab,tenant_id=acme, promptHi team, how do I reset my password? Thanks in advance!→"hit": true, "stage": "exact". The greeting and sign-off were normalized away onto the same key.How can I reset my password?at threshold0.80→"stage": "cosine"(measured similarity ≈ 0.85; at the default0.92it misses and returnsbest_similarity0.85 — read that number before choosing a threshold).- Wrong-hit protection:
Why do I need to reset my password?at threshold0.70→"reason": "guard_rejected"withquestion_type_mismatch:q={why},c={how}. High similarity is not enough. - Tenant isolation: switch
tenant_idtoglobexand repeat →"reason": "empty_index". Another tenant's entries are unreachable.
Dashboard → Refresh: one-line summary (requests / hits / misses / hit rate / saved tokens / estimated saving), hit-rate time series, cumulative savings per tenant, p90 latency, most-reused entries, index health. Export events CSV takes the log with you (prompt text is not recorded by default).
Sweep → Download a sample log, upload that same file, enter your prices, Run sweep. You get "at this threshold you reuse X% and save $Y", with MismatchGuard on and off side by side — the conservative row is the one you can defend in a review.
API Docs tab shows the real paths for your URL and your Gradio version. The Use via API link in the page footer is generated by Gradio itself and is always authoritative.
From the command line
SPACE="https://<user>-echocache.hf.space"
EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/store" -H "Content-Type: application/json" \
-d '{"data": ["acme","How do I reset my password?","Open Settings > Security > Reset password.","support",86400]}' \
| python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])")
curl -s -N "$SPACE/gradio_api/call/store/$EVENT_ID"
EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/lookup" -H "Content-Type: application/json" \
-d '{"data": ["acme","Hi team, how do I reset my password? Thanks!",0.92,"support"]}' \
| python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])")
curl -s -N "$SPACE/gradio_api/call/lookup/$EVENT_ID"
pip install gradio_client
python3 - <<'PY'
from gradio_client import Client
c = Client("<user>/EchoCache")
print(c.predict("acme", "how do i reset my password", 0.92, "support", api_name="/lookup"))
print(c.predict(api_name="/health"))
PY
5. How to run it
5-1. Export before the Space sleeps, import after it wakes
The free tier has no persistent disk and sleeps after 48h of inactivity (the timer cannot be changed). The index lives in RAM only, so carrying it over is an explicit step:
Backuptab → tick include response bodies → Export index- Save
index.jsonsomewhere outside the Space - After a sleep or restart:
Backuptab → uploadindex.json→ Import index - Check
imported / skipped_no_response / expired / rejected_by_safetyin the result
- Vectors are recomputed on import, so a changed
VECTOR_DIMis not a problem. - The safety screen runs again on import (leave
skip_safetyoff). - An export without response bodies is for analysis only — those rows are skipped on import and counted under
skipped_no_response.
Daily backup from your own machine:
python3 - <<'PY'
from gradio_client import Client
import datetime
c = Client("<user>/EchoCache")
path, receipt = c.predict(True, "", api_name="/export_index")
dst = "echocache-%s.json" % datetime.date.today()
open(dst, "w").write(open(path).read())
print(dst, receipt)
PY
5-2. Choose the threshold from a Sweep, not from intuition
- The default
0.92is conservative: it mostly reuses exact matches and surface variations (greetings, honorifics, punctuation, casing). - Genuine paraphrases with different wording land around 0.80–0.90 with character n-grams.
- Export a day of real traffic as CSV (
tenant_id, prompt, response, route), run Sweep, and read reuse rate, estimated saving, borderline count and guard blocks before deciding. - Before lowering the threshold, open Audit and read why the guard rejected candidates. Many rejections mean "these questions genuinely differ", not "the threshold is too high".
- On the published benchmark, threshold 0.80 with the guard on scored 99.2% accuracy with 100% reuse recall and 1 wrong reuse in 90 — but that is one synthetic set, not your traffic. Measure yours.
5-3. What to watch
| Frequency | Where | What you are deciding |
|---|---|---|
| daily | Dashboard hit rate / saved cost | is reuse growing as expected |
| daily | Audit → guard rejections | rejection patterns; early warning of wrong hits |
| weekly | Dashboard index health (evicted / expired) |
revisit TTL or the entry caps |
| weekly | Sweep on fresh logs | revisit the threshold |
| on change | Health → mode |
did it silently switch away from local-only |
6. Troubleshooting
Memory pressure, or the Space restarts
Index memory is roughly entries × VECTOR_DIM × 4 bytes, and MAX_ENTRIES is per tenant. With the defaults, a full tenant is ≈ 80 MB and the whole index is capped at ≈ 330 MB by MAX_TOTAL_ENTRIES=20000.
In order of effect (Settings → Variables and secrets):
MAX_TOTAL_ENTRIES = 8000 # the global ceiling — start here
MAX_ENTRIES = 2000 # per tenant
MAX_TEXT_CHARS = 4000 # per stored prompt / response
VECTOR_DIM = 2048 # halves memory, costs little accuracy
EVENT_BUFFER = 2000 # event ring buffer
Health shows the live approx_memory_mb and vector_matrix_mb_per_full_tenant.
A tenant's cache went empty
Reaching MAX_TENANTS (default 256) evicts the least-recently-used tenant's entire partition on the next new tenant's first write. Deletion only — nothing is ever moved or read across tenants. Size MAX_TENANTS above your real tenant count. When the global entry budget is reached instead, entries are evicted one at a time from the largest partition.
Nothing hits
Work through this in order before touching the threshold:
- Read
reasonin the Lookup result:empty_index— nothing stored for that tenant (a mistypedtenant_idis the usual cause)no_candidates— no fingerprint-similar entry; the wording differs a lotbelow_threshold—best_similarityis in the payload; decide from that numberguard_rejected— the meaning was judged different; go to step 2
- Open Audit and read the reasons (
numeric_mismatch,negation_mismatch,proper_noun_mismatch,temporal_mismatch,question_type_mismatch). If the reason is sound, that miss is correct and lowering the threshold will not fix it. - Compare the stored and queried prompts:
cached_prompt_previewcomes back in the response. - Still missing legitimate reuse? Step down through Sweep (0.86 → 0.82) and watch
borderline_hitsandguard_blocksgrow. - Only if a specific check genuinely does not fit your domain, disable it individually:
GUARD_DISABLE=proper_noun. Disabling all of them is not recommended.
A wrong hit happened (deal with this first)
- Raise the threshold (e.g.
0.92→0.95). Immediate effect. - Find the
entry_idin Audit, or in Dashboard → most reused entries. - Use Invalidate entry at the bottom of the Audit tab (
tenant_id+ the first 16 characters of the id). A route name invalidates that route;*clears the tenant. - If the same shape recurs, add those prompts to a CSV, re-run Sweep, and adopt a threshold that rejects them.
- If the difference is not one the five checks can see, stop caching that category: give it its own
routeand invalidate that route on a schedule.
The embedding API returns 402 or 429
402 Payment Required— the free inference credit (~$0.10/month) is exhausted.429 Too Many Requests— rate limited.
Neither stops the cache. The exception is swallowed and matching continues on local vectors alone. After repeated failures a circuit breaker opens for EMBED_COOLDOWN_SEC (default 600s) so no further credit is spent. Health shows embedding.last_error and cooldown_remaining_sec. To stop permanently, delete HF_TOKEN from Secrets and restart — the header returns to local-only. If it does not, check for HUGGINGFACE_HUB_TOKEN / HUGGINGFACEHUB_API_TOKEN.
The Space sleeps after 48h and wakes up empty
That is the free tier working as designed (no persistent disk, fixed sleep timer). Use the export/import routine in 5-1. If you need it always on, move to paid hardware or run python app.py on your own host — the same file works unchanged.
/store returns stored: false
The safety screen refused the content. Read reason:
| reason | meaning |
|---|---|
credit_card_luhn |
a digit run passed the Luhn checksum; order numbers are not flagged |
credential_prefix |
sk-, ghp_, AKIA, xoxb-, AIza, hf_, … |
jwt_structure |
a three-part JWT |
high_entropy_secret |
a long high-entropy run that is not identifier-shaped |
contact_pii_excess |
too many e-mail addresses or phone numbers |
private_key_block |
a PEM private-key block |
This is correct behaviour. Before relaxing anything, ask whether that content belongs in a cache at all.
The API path 404s
Gradio changes it between major versions (5 and 6 use /gradio_api/call/<name>, 4 uses /call/<name>). Do not hardcode it — read it from the API Docs tab or the Use via API footer link. gradio_client resolves it for you.
Appendix: running locally
pip install -r requirements.txt
python3 app.py # http://localhost:7860
Identical behaviour to the Space. Environment variables work the same way:
PRICE_IN_PER_1K=0.003 PRICE_OUT_PER_1K=0.015 MAX_ENTRIES=1000 python3 app.py