File size: 2,804 Bytes
49a8dc7
4123863
49a8dc7
4123863
49a8dc7
4123863
49a8dc7
 
 
 
4123863
49a8dc7
4123863
49a8dc7
 
 
 
 
 
4123863
49a8dc7
 
 
 
 
 
 
 
 
4123863
 
 
 
49a8dc7
4123863
 
 
49a8dc7
4123863
49a8dc7
4123863
49a8dc7
4123863
49a8dc7
4123863
49a8dc7
 
 
 
 
 
 
 
 
 
 
 
4123863
 
 
 
49a8dc7
4123863
 
49a8dc7
4123863
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
# Deploying to LukeFP/Physh_Classification

The Space repo lives at `~/code/2026.7/Physh_Classification`.

## 1. Add the token secret

`google/embeddinggemma-300m` is gated. Accept the Gemma license while signed in,
create a **read** token, then on the Space page: Settings β†’ *Variables and
secrets* β†’ **New secret**, name `HF_TOKEN`, value the token. Without it the Space
boots fine but the first classification fails with a 401.

## 2. Hardware

On the free tier, Gradio Spaces run on **ZeroGPU**, which stops the container at
startup unless it finds at least one `@spaces.GPU` function β€” the
`No @spaces.GPU function detected during startup` error. `infer()` in `app.py`
carries that decorator, so ZeroGPU is satisfied.

Constraints ZeroGPU imposes, and how `app.py` meets them:

| Constraint | Handling |
|---|---|
| `import spaces` must precede `import torch` | It is the first import in `app.py` |
| Nothing may touch CUDA outside a `@GPU` function | Models load with `device="cpu"`; `.to(device)` happens inside `infer()` |
| Return values cross a process boundary | `infer()` returns plain `list[float]`, never CUDA tensors |
| One GPU allocation per call, with a duration budget | `@GPU(duration=60)`; the model is already resident, so only the encode runs |

CPU basic (a PRO perk) also works with this code unchanged β€” `spaces` is an
optional import and the device is chosen from `torch.cuda.is_available()`.

## 3. Push

```bash
cd ~/code/2026.7/Physh_Classification
git push origin main
```

The build takes a few minutes, most of it `pip install torch`.

## 4. First checks

- **Predictions look like noise, or nothing clears the threshold.** Almost
  certainly the embedding prompt. Open *Advanced* and try the other two formats;
  the one matching your training pipeline gives confident, coherent labels.
  Once you know which, set `DEFAULT_PROMPT` at the top of `app.py`.
  (`~/code/2026/embedding_title_abstract` likely has the answer.)
- **Error mentioning a gated repo, or a 401.** `HF_TOKEN` is missing, wrong, or
  the account behind it hasn't accepted the Gemma license.
- **First request is slow, later ones fast.** Expected β€” EmbeddingGemma loads
  lazily on first use so the Space boots quickly. Cached after that.

## Updating later

Retraining only needs a push to
[`LukeFP/physh_topic_supervised_classifier`](https://huggingface.co/LukeFP/physh_topic_supervised_classifier);
the Space picks up new weights on its next restart. Only change this repo if the
*filenames* change β€” they're the constants at the top of `app.py`.

## Local smoke test

Runs the real checkpoints through the full chain with a stubbed embedder, so it
needs no token and no model download:

```bash
PHYSH_WEIGHTS_DIR=~/code/2026.7/physh_topic_supervised_classifier python test_local.py
```