LukeFP's picture
Run on ZeroGPU: add @spaces.GPU entry point
49a8dc7
|
Raw History Blame Contribute Delete
2.8 kB
# Deploying to LukeFP/Physh_Classification
The Space repo lives at `~/code/2026.7/Physh_Classification`.
## 1. Add the token secret
`google/embeddinggemma-300m` is gated. Accept the Gemma license while signed in,
create a **read** token, then on the Space page: Settings β†’ *Variables and
secrets* β†’ **New secret**, name `HF_TOKEN`, value the token. Without it the Space
boots fine but the first classification fails with a 401.
## 2. Hardware
On the free tier, Gradio Spaces run on **ZeroGPU**, which stops the container at
startup unless it finds at least one `@spaces.GPU` function β€” the
`No @spaces.GPU function detected during startup` error. `infer()` in `app.py`
carries that decorator, so ZeroGPU is satisfied.
Constraints ZeroGPU imposes, and how `app.py` meets them:
| Constraint | Handling |
|---|---|
| `import spaces` must precede `import torch` | It is the first import in `app.py` |
| Nothing may touch CUDA outside a `@GPU` function | Models load with `device="cpu"`; `.to(device)` happens inside `infer()` |
| Return values cross a process boundary | `infer()` returns plain `list[float]`, never CUDA tensors |
| One GPU allocation per call, with a duration budget | `@GPU(duration=60)`; the model is already resident, so only the encode runs |
CPU basic (a PRO perk) also works with this code unchanged β€” `spaces` is an
optional import and the device is chosen from `torch.cuda.is_available()`.
## 3. Push
```bash
cd ~/code/2026.7/Physh_Classification
git push origin main
```
The build takes a few minutes, most of it `pip install torch`.
## 4. First checks
- **Predictions look like noise, or nothing clears the threshold.** Almost
certainly the embedding prompt. Open *Advanced* and try the other two formats;
the one matching your training pipeline gives confident, coherent labels.
Once you know which, set `DEFAULT_PROMPT` at the top of `app.py`.
(`~/code/2026/embedding_title_abstract` likely has the answer.)
- **Error mentioning a gated repo, or a 401.** `HF_TOKEN` is missing, wrong, or
the account behind it hasn't accepted the Gemma license.
- **First request is slow, later ones fast.** Expected β€” EmbeddingGemma loads
lazily on first use so the Space boots quickly. Cached after that.
## Updating later
Retraining only needs a push to
[`LukeFP/physh_topic_supervised_classifier`](https://huggingface.co/LukeFP/physh_topic_supervised_classifier);
the Space picks up new weights on its next restart. Only change this repo if the
*filenames* change β€” they're the constants at the top of `app.py`.
## Local smoke test
Runs the real checkpoints through the full chain with a stubbed embedder, so it
needs no token and no model download:
```bash
PHYSH_WEIGHTS_DIR=~/code/2026.7/physh_topic_supervised_classifier python test_local.py
```