File size: 13,979 Bytes
46985e0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
# Deploying EchoCache on a free Hugging Face Space

> 日本語版: [`deploy.ja.md`](deploy.ja.md)

About 10 minutes, copy-paste throughout. You need a browser, `git` and `python3`.
No paid plan, no GPU, no model download.

---

## 1. Create the Space

1. Open <https://huggingface.co/new-space>
2. Fill in:

   | Field | Value |
   |---|---|
   | **Owner** | your account |
   | **Space name** | `EchoCache` (referred to below as `<user>/EchoCache`) |
   | **License** | `apache-2.0` |
   | **SDK** | **Gradio** |
   | **Space hardware** | **CPU basic · 2 vCPU · 16 GB · FREE** |
   | **Visibility** | Public or Private — both work |

3. Press **Create Space**. If a Gradio template picker appears, choose "Blank"; every file gets replaced anyway.

---

## 2. Push the files

Log in once with a **write** token from <https://huggingface.co/settings/tokens>:

```bash
hf auth login          # older installs: huggingface-cli login
```

Then run this as-is, replacing `<user>` with your account name:

```bash
USER=<user>; SPACE=EchoCache; SRC="$(pwd)"
git clone https://huggingface.co/spaces/$USER/$SPACE ~/$SPACE-space && \
cp "$SRC/app.py" "$SRC/requirements.txt" "$SRC/deploy.md" "$SRC/deploy.ja.md" ~/$SPACE-space/ && \
cp "$SRC/README.gradio.md" ~/$SPACE-space/README.md && \
cd ~/$SPACE-space && git add app.py requirements.txt README.md deploy.md deploy.ja.md && \
git commit -m "Deploy EchoCache: semantic cache + cost observability" && git push
```

If `git clone` asks for credentials, use your account name as the username and your **access token** as the password.

The Space builds automatically (**Building** → **Running**, 1–2 minutes on a first push). Watch **Logs** on the Space page.

> Note the card swap: the published `README.md` carries `sdk: static`, so a Gradio deployment must use `README.gradio.md`.

> If the build complains about `sdk_version`, change that line in `README.md` to a version offered in your Space's Settings, then push again. Nothing else needs to change.

---

## 3. Set `HF_TOKEN` (optional — everything works without it)

Only needed if you want the optional embedding re-rank.

1. Space page → **Settings** → **Variables and secrets**
2. **New secret** → Name: `HF_TOKEN` / Value: your token → **Save**
   (`HUGGINGFACE_HUB_TOKEN` and `HUGGINGFACEHUB_API_TOKEN` are read as well. If one of them is exported in your shell during a local run, you will start in `local+embedding` mode and make outbound calls — the header and the `Health` tab always say which mode is active.)
3. **Restart this Space**

| | without `HF_TOKEN` | with `HF_TOKEN` |
|---|---|---|
| similarity | fully local (default, recommended) | local + embedding, weighted average |
| core features | **all working** | all working |
| on failure | — | 402/429/timeout falls back to local automatically |

A free Hugging Face account has roughly **$0.10/month** of inference credit. **Running without the token is the normal configuration.**

Useful optional variables, same screen:

```
PRICE_IN_PER_1K   = 0.003     # your input price; leave unset and savings stay $0
PRICE_OUT_PER_1K  = 0.015     # your output price
MAX_ENTRIES       = 5000      # per tenant
MAX_TOTAL_ENTRIES = 20000     # across all tenants — the real memory guard
MAX_TENANTS       = 256       # partition cap; exceeding it drops the LRU tenant's whole partition
DEFAULT_THRESHOLD = 0.92      # conservative by default; move it after running a Sweep
```

---

## 4. Verify it works (3 minutes in the UI)

1. **Open the Space** (`https://<user>-echocache.hf.space`). The header should read `Mode: local-only`.

2. **Store one entry** — `Store` tab, `tenant_id` = `acme`, `route` = `support`,
   prompt `How do I reset my password?`, response `Open Settings > Security > Reset password.` → **Store** returns `"stored": true` and an `entry_id`.
   (Or press *Seed 3 demo entries*.)

3. **Look it up**
   * `Lookup` tab, `tenant_id` = `acme`, prompt `Hi team, how do I reset my password? Thanks in advance!` → `"hit": true, "stage": "exact"`. The greeting and sign-off were normalized away onto the same key.
   * `How can I reset my password?` at threshold `0.80` → `"stage": "cosine"` (measured similarity ≈ **0.85**; at the default `0.92` it misses and returns `best_similarity` 0.85 — **read that number before choosing a threshold**).
   * **Wrong-hit protection:** `Why do I need to reset my password?` at threshold `0.70` → `"reason": "guard_rejected"` with `question_type_mismatch:q={why},c={how}`. High similarity is not enough.
   * **Tenant isolation:** switch `tenant_id` to `globex` and repeat → `"reason": "empty_index"`. Another tenant's entries are unreachable.

4. **Dashboard** → **Refresh**: one-line summary (requests / hits / misses / hit rate / saved tokens / estimated saving), hit-rate time series, cumulative savings per tenant, p90 latency, most-reused entries, index health. **Export events CSV** takes the log with you (prompt text is not recorded by default).

5. **Sweep** → **Download a sample log**, upload that same file, enter your prices, **Run sweep**. You get "at this threshold you reuse X% and save $Y", with MismatchGuard **on** and **off** side by side — the conservative row is the one you can defend in a review.

6. **API Docs** tab shows the real paths for your URL and your Gradio version. The **Use via API** link in the page footer is generated by Gradio itself and is always authoritative.

### From the command line

```bash
SPACE="https://<user>-echocache.hf.space"

EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/store" -H "Content-Type: application/json" \
  -d '{"data": ["acme","How do I reset my password?","Open Settings > Security > Reset password.","support",86400]}' \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])")
curl -s -N "$SPACE/gradio_api/call/store/$EVENT_ID"

EVENT_ID=$(curl -s -X POST "$SPACE/gradio_api/call/lookup" -H "Content-Type: application/json" \
  -d '{"data": ["acme","Hi team, how do I reset my password? Thanks!",0.92,"support"]}' \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['event_id'])")
curl -s -N "$SPACE/gradio_api/call/lookup/$EVENT_ID"
```

```bash
pip install gradio_client
python3 - <<'PY'
from gradio_client import Client
c = Client("<user>/EchoCache")
print(c.predict("acme", "how do i reset my password", 0.92, "support", api_name="/lookup"))
print(c.predict(api_name="/health"))
PY
```

---

## 5. How to run it

### 5-1. Export before the Space sleeps, import after it wakes

The free tier has **no persistent disk** and **sleeps after 48h of inactivity** (the timer cannot be changed). The index lives in RAM only, so carrying it over is an explicit step:

1. `Backup` tab → tick *include response bodies* → **Export index**
2. Save `index.json` somewhere outside the Space
3. After a sleep or restart: `Backup` tab → upload `index.json` → **Import index**
4. Check `imported / skipped_no_response / expired / rejected_by_safety` in the result

* Vectors are **recomputed** on import, so a changed `VECTOR_DIM` is not a problem.
* The safety screen **runs again** on import (leave `skip_safety` off).
* An export without response bodies is for analysis only — those rows are skipped on import and counted under `skipped_no_response`.

Daily backup from your own machine:

```bash
python3 - <<'PY'
from gradio_client import Client
import datetime
c = Client("<user>/EchoCache")
path, receipt = c.predict(True, "", api_name="/export_index")
dst = "echocache-%s.json" % datetime.date.today()
open(dst, "w").write(open(path).read())
print(dst, receipt)
PY
```

### 5-2. Choose the threshold from a Sweep, not from intuition

* The default `0.92` is **conservative**: it mostly reuses exact matches and surface variations (greetings, honorifics, punctuation, casing).
* Genuine paraphrases with different wording land around **0.80–0.90** with character n-grams.
* Export a day of real traffic as CSV (`tenant_id, prompt, response, route`), run **Sweep**, and read reuse rate, estimated saving, borderline count and guard blocks before deciding.
* Before lowering the threshold, open **Audit** and read why the guard rejected candidates. Many rejections mean "these questions genuinely differ", not "the threshold is too high".
* On the published benchmark, threshold **0.80 with the guard on** scored 99.2% accuracy with 100% reuse recall and 1 wrong reuse in 90 — but that is one synthetic set, not your traffic. Measure yours.

### 5-3. What to watch

| Frequency | Where | What you are deciding |
|---|---|---|
| daily | Dashboard hit rate / saved cost | is reuse growing as expected |
| daily | Audit → guard rejections | rejection patterns; early warning of wrong hits |
| weekly | Dashboard index health (`evicted` / `expired`) | revisit TTL or the entry caps |
| weekly | Sweep on fresh logs | revisit the threshold |
| on change | Health → `mode` | did it silently switch away from `local-only` |

---

## 6. Troubleshooting

### Memory pressure, or the Space restarts

Index memory is roughly `entries × VECTOR_DIM × 4 bytes`, and **`MAX_ENTRIES` is per tenant**. With the defaults, a full tenant is ≈ 80 MB and the whole index is capped at ≈ 330 MB by `MAX_TOTAL_ENTRIES=20000`.

In order of effect (Settings → Variables and secrets):

```
MAX_TOTAL_ENTRIES = 8000      # the global ceiling — start here
MAX_ENTRIES       = 2000      # per tenant
MAX_TEXT_CHARS    = 4000      # per stored prompt / response
VECTOR_DIM        = 2048      # halves memory, costs little accuracy
EVENT_BUFFER      = 2000      # event ring buffer
```

`Health` shows the live `approx_memory_mb` and `vector_matrix_mb_per_full_tenant`.

### A tenant's cache went empty

Reaching `MAX_TENANTS` (default 256) **evicts the least-recently-used tenant's entire partition** on the next new tenant's first write. Deletion only — nothing is ever moved or read across tenants. Size `MAX_TENANTS` above your real tenant count. When the global entry budget is reached instead, entries are evicted one at a time from the largest partition.

### Nothing hits

Work through this in order before touching the threshold:

1. Read `reason` in the Lookup result:
   * `empty_index` — nothing stored for that tenant (a mistyped `tenant_id` is the usual cause)
   * `no_candidates` — no fingerprint-similar entry; the wording differs a lot
   * `below_threshold` — `best_similarity` is in the payload; decide from that number
   * `guard_rejected` — the meaning was judged different; go to step 2
2. Open **Audit** and read the reasons (`numeric_mismatch`, `negation_mismatch`, `proper_noun_mismatch`, `temporal_mismatch`, `question_type_mismatch`). If the reason is sound, that miss is **correct** and lowering the threshold will not fix it.
3. Compare the stored and queried prompts: `cached_prompt_preview` comes back in the response.
4. Still missing legitimate reuse? Step down through Sweep (0.86 → 0.82) and watch `borderline_hits` and `guard_blocks` grow.
5. Only if a specific check genuinely does not fit your domain, disable it individually: `GUARD_DISABLE=proper_noun`. Disabling all of them is not recommended.

### A wrong hit happened (deal with this first)

1. **Raise the threshold** (e.g. `0.92` → `0.95`). Immediate effect.
2. Find the `entry_id` in **Audit**, or in Dashboard → *most reused entries*.
3. Use **Invalidate entry** at the bottom of the Audit tab (`tenant_id` + the first 16 characters of the id). A route name invalidates that route; `*` clears the tenant.
4. If the same shape recurs, add those prompts to a CSV, re-run Sweep, and adopt a threshold that rejects them.
5. If the difference is not one the five checks can see, stop caching that category: give it its own `route` and invalidate that route on a schedule.

### The embedding API returns 402 or 429

* `402 Payment Required` — the free inference credit (~$0.10/month) is exhausted.
* `429 Too Many Requests` — rate limited.

**Neither stops the cache.** The exception is swallowed and matching continues on local vectors alone. After repeated failures a circuit breaker opens for `EMBED_COOLDOWN_SEC` (default 600s) so no further credit is spent. `Health` shows `embedding.last_error` and `cooldown_remaining_sec`. To stop permanently, delete `HF_TOKEN` from Secrets and restart — the header returns to `local-only`. If it does not, check for `HUGGINGFACE_HUB_TOKEN` / `HUGGINGFACEHUB_API_TOKEN`.

### The Space sleeps after 48h and wakes up empty

That is the free tier working as designed (no persistent disk, fixed sleep timer). Use the export/import routine in 5-1. If you need it always on, move to paid hardware or run `python app.py` on your own host — the same file works unchanged.

### `/store` returns `stored: false`

The safety screen refused the content. Read `reason`:

| reason | meaning |
|---|---|
| `credit_card_luhn` | a digit run passed the Luhn checksum; order numbers are not flagged |
| `credential_prefix` | `sk-`, `ghp_`, `AKIA`, `xoxb-`, `AIza`, `hf_`, … |
| `jwt_structure` | a three-part JWT |
| `high_entropy_secret` | a long high-entropy run that is not identifier-shaped |
| `contact_pii_excess` | too many e-mail addresses or phone numbers |
| `private_key_block` | a PEM private-key block |

**This is correct behaviour.** Before relaxing anything, ask whether that content belongs in a cache at all.

### The API path 404s

Gradio changes it between major versions (5 and 6 use `/gradio_api/call/<name>`, 4 uses `/call/<name>`). Do not hardcode it — read it from the **API Docs** tab or the **Use via API** footer link. `gradio_client` resolves it for you.

---

## Appendix: running locally

```bash
pip install -r requirements.txt
python3 app.py          # http://localhost:7860
```

Identical behaviour to the Space. Environment variables work the same way:

```bash
PRICE_IN_PER_1K=0.003 PRICE_OUT_PER_1K=0.015 MAX_ENTRIES=1000 python3 app.py
```