Spaces:
Running on Zero
Running on Zero
File size: 2,804 Bytes
49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 49a8dc7 4123863 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 | # Deploying to LukeFP/Physh_Classification
The Space repo lives at `~/code/2026.7/Physh_Classification`.
## 1. Add the token secret
`google/embeddinggemma-300m` is gated. Accept the Gemma license while signed in,
create a **read** token, then on the Space page: Settings β *Variables and
secrets* β **New secret**, name `HF_TOKEN`, value the token. Without it the Space
boots fine but the first classification fails with a 401.
## 2. Hardware
On the free tier, Gradio Spaces run on **ZeroGPU**, which stops the container at
startup unless it finds at least one `@spaces.GPU` function β the
`No @spaces.GPU function detected during startup` error. `infer()` in `app.py`
carries that decorator, so ZeroGPU is satisfied.
Constraints ZeroGPU imposes, and how `app.py` meets them:
| Constraint | Handling |
|---|---|
| `import spaces` must precede `import torch` | It is the first import in `app.py` |
| Nothing may touch CUDA outside a `@GPU` function | Models load with `device="cpu"`; `.to(device)` happens inside `infer()` |
| Return values cross a process boundary | `infer()` returns plain `list[float]`, never CUDA tensors |
| One GPU allocation per call, with a duration budget | `@GPU(duration=60)`; the model is already resident, so only the encode runs |
CPU basic (a PRO perk) also works with this code unchanged β `spaces` is an
optional import and the device is chosen from `torch.cuda.is_available()`.
## 3. Push
```bash
cd ~/code/2026.7/Physh_Classification
git push origin main
```
The build takes a few minutes, most of it `pip install torch`.
## 4. First checks
- **Predictions look like noise, or nothing clears the threshold.** Almost
certainly the embedding prompt. Open *Advanced* and try the other two formats;
the one matching your training pipeline gives confident, coherent labels.
Once you know which, set `DEFAULT_PROMPT` at the top of `app.py`.
(`~/code/2026/embedding_title_abstract` likely has the answer.)
- **Error mentioning a gated repo, or a 401.** `HF_TOKEN` is missing, wrong, or
the account behind it hasn't accepted the Gemma license.
- **First request is slow, later ones fast.** Expected β EmbeddingGemma loads
lazily on first use so the Space boots quickly. Cached after that.
## Updating later
Retraining only needs a push to
[`LukeFP/physh_topic_supervised_classifier`](https://huggingface.co/LukeFP/physh_topic_supervised_classifier);
the Space picks up new weights on its next restart. Only change this repo if the
*filenames* change β they're the constants at the top of `app.py`.
## Local smoke test
Runs the real checkpoints through the full chain with a stubbed embedder, so it
needs no token and no model download:
```bash
PHYSH_WEIGHTS_DIR=~/code/2026.7/physh_topic_supervised_classifier python test_local.py
```
|