ukung's picture
Add source code (encoder_loader, model, data, train, ablation, generate) + NOTES
fd3090c verified
|
Raw History Blame Contribute Delete
3.08 kB
# NOTES β€” gotchas and current state
Everything below was hit for real while building this. Read it before changing
anything, especially the encoder loading.
## Environment
**`transformers` must be pinned to 4.57.1.** Version 5.x breaks the
Semantic-Lite-2 custom code in three separate ways (`create_causal_mask`
signature, the rope validator, tied-weight handling). `requirements.txt` pins it.
## Encoder loading (the tricky part)
`ukung/semantic-lite-2` uses custom modelling code, so `AutoModel.from_pretrained`
does not work. Two source patches are required, both applied at load time in
`encoder_loader.py`:
1. `"inputs_embeds"` β†’ `"input_embeds"` β€” the keyword `create_causal_mask`
expects in 4.57.1.
2. Add `"cache_position": cache_position` to `mask_kwargs` β€” required in 4.57.1.
Also: `rope_parameters=None` is passed to the config to bypass the validator,
then set manually afterwards. Without this the config constructor rejects the
rope settings.
After patching, `sys.modules` entries starting with `semantic_lite` are purged
so the patched file is what actually gets imported. Skipping that step silently
loads a stale module.
## Frozen encoder must stay in eval()
`model.train()` recurses into every submodule β€” including the frozen encoder,
whose attention heads carry `dropout=0.1`. That injects noise into Data A and
Data B on every step. Always use `set_train_mode(model)` from
`encoder_loader.py`, never a bare `.train()`.
## Token alignment
```
tgt_ids = [t0, t1, ..., tN, EOS] <- no BOS
dec_input = [BOS, t0, t1, ..., tN] <- shifted right
labels = [t0, t1, ..., tN, EOS]
```
EOS lands in `dec_input` at a position whose label is PAD. That is correct (the
token before the first PAD is EOS) and no gradient flows there.
## Memory
Logits are `(B, L, 131072)`. At L=768 that is ~100M floats per sample, so
`batch_size=1` is mandatory on a T4 β€” batch>1 OOMs. Gradient accumulation (4)
recovers an effective batch of 4.
## Results so far (smoke test, 1000 samples, 10 epochs, T4)
| Metric | Value |
|---|---|
| Train loss (final) | 4.6753 |
| Eval loss (final) | 5.3534 |
**The model has not learned reasoning.** Generated text reproduces the surface
format of the training data ("Let me analyze this problem step by step") but the
content is not grounded in the input.
## Open questions
1. **Is the conditioning doing anything?** `ablation.py` has not been run. Until
it is, the contribution of Data A + Data B is unverified. This is the single
most important experiment for this architecture.
2. **Tied embedding.** The output head reuses the encoder's frozen embedding
matrix, saving ~132M params. But that matrix was never trained for
generation, so the output geometry may be a poor fit. An untied head is worth
trying.
3. **Data B source.** Currently the last hidden state of the 2-layer backbone.
Intermediate layers, or the attention-head layers, may carry better signal.
4. **Scale.** 1000 samples / 10 epochs is a smoke test, not a training run.
Eval loss 5.35 is far from usable.