File size: 3,076 Bytes
fd3090c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 | # NOTES — gotchas and current state
Everything below was hit for real while building this. Read it before changing
anything, especially the encoder loading.
## Environment
**`transformers` must be pinned to 4.57.1.** Version 5.x breaks the
Semantic-Lite-2 custom code in three separate ways (`create_causal_mask`
signature, the rope validator, tied-weight handling). `requirements.txt` pins it.
## Encoder loading (the tricky part)
`ukung/semantic-lite-2` uses custom modelling code, so `AutoModel.from_pretrained`
does not work. Two source patches are required, both applied at load time in
`encoder_loader.py`:
1. `"inputs_embeds"` → `"input_embeds"` — the keyword `create_causal_mask`
expects in 4.57.1.
2. Add `"cache_position": cache_position` to `mask_kwargs` — required in 4.57.1.
Also: `rope_parameters=None` is passed to the config to bypass the validator,
then set manually afterwards. Without this the config constructor rejects the
rope settings.
After patching, `sys.modules` entries starting with `semantic_lite` are purged
so the patched file is what actually gets imported. Skipping that step silently
loads a stale module.
## Frozen encoder must stay in eval()
`model.train()` recurses into every submodule — including the frozen encoder,
whose attention heads carry `dropout=0.1`. That injects noise into Data A and
Data B on every step. Always use `set_train_mode(model)` from
`encoder_loader.py`, never a bare `.train()`.
## Token alignment
```
tgt_ids = [t0, t1, ..., tN, EOS] <- no BOS
dec_input = [BOS, t0, t1, ..., tN] <- shifted right
labels = [t0, t1, ..., tN, EOS]
```
EOS lands in `dec_input` at a position whose label is PAD. That is correct (the
token before the first PAD is EOS) and no gradient flows there.
## Memory
Logits are `(B, L, 131072)`. At L=768 that is ~100M floats per sample, so
`batch_size=1` is mandatory on a T4 — batch>1 OOMs. Gradient accumulation (4)
recovers an effective batch of 4.
## Results so far (smoke test, 1000 samples, 10 epochs, T4)
| Metric | Value |
|---|---|
| Train loss (final) | 4.6753 |
| Eval loss (final) | 5.3534 |
**The model has not learned reasoning.** Generated text reproduces the surface
format of the training data ("Let me analyze this problem step by step") but the
content is not grounded in the input.
## Open questions
1. **Is the conditioning doing anything?** `ablation.py` has not been run. Until
it is, the contribution of Data A + Data B is unverified. This is the single
most important experiment for this architecture.
2. **Tied embedding.** The output head reuses the encoder's frozen embedding
matrix, saving ~132M params. But that matrix was never trained for
generation, so the output geometry may be a poor fit. An untied head is worth
trying.
3. **Data B source.** Currently the last hidden state of the 2-layer backbone.
Intermediate layers, or the attention-head layers, may carry better signal.
4. **Scale.** 1000 samples / 10 epochs is a smoke test, not a training run.
Eval loss 5.35 is far from usable.
|