|
Download code/NOTES.md from ukung/semantic-lite-2-decoder-smoke-test: direct link, hf CLI and curl.
- Browser
- Download file 3.08 kB
-
https://huggingface.co/ukung/semantic-lite-2-decoder-smoke-test/resolve/main/code/NOTES.md
- Command line
-
hf download hf://ukung/semantic-lite-2-decoder-smoke-test/code/NOTES.md
-
curl -L -o NOTES.md https://huggingface.co/ukung/semantic-lite-2-decoder-smoke-test/resolve/main/code/NOTES.md
3.08 kB
| # NOTES β gotchas and current state | |
| Everything below was hit for real while building this. Read it before changing | |
| anything, especially the encoder loading. | |
| ## Environment | |
| **`transformers` must be pinned to 4.57.1.** Version 5.x breaks the | |
| Semantic-Lite-2 custom code in three separate ways (`create_causal_mask` | |
| signature, the rope validator, tied-weight handling). `requirements.txt` pins it. | |
| ## Encoder loading (the tricky part) | |
| `ukung/semantic-lite-2` uses custom modelling code, so `AutoModel.from_pretrained` | |
| does not work. Two source patches are required, both applied at load time in | |
| `encoder_loader.py`: | |
| 1. `"inputs_embeds"` β `"input_embeds"` β the keyword `create_causal_mask` | |
| expects in 4.57.1. | |
| 2. Add `"cache_position": cache_position` to `mask_kwargs` β required in 4.57.1. | |
| Also: `rope_parameters=None` is passed to the config to bypass the validator, | |
| then set manually afterwards. Without this the config constructor rejects the | |
| rope settings. | |
| After patching, `sys.modules` entries starting with `semantic_lite` are purged | |
| so the patched file is what actually gets imported. Skipping that step silently | |
| loads a stale module. | |
| ## Frozen encoder must stay in eval() | |
| `model.train()` recurses into every submodule β including the frozen encoder, | |
| whose attention heads carry `dropout=0.1`. That injects noise into Data A and | |
| Data B on every step. Always use `set_train_mode(model)` from | |
| `encoder_loader.py`, never a bare `.train()`. | |
| ## Token alignment | |
| ``` | |
| tgt_ids = [t0, t1, ..., tN, EOS] <- no BOS | |
| dec_input = [BOS, t0, t1, ..., tN] <- shifted right | |
| labels = [t0, t1, ..., tN, EOS] | |
| ``` | |
| EOS lands in `dec_input` at a position whose label is PAD. That is correct (the | |
| token before the first PAD is EOS) and no gradient flows there. | |
| ## Memory | |
| Logits are `(B, L, 131072)`. At L=768 that is ~100M floats per sample, so | |
| `batch_size=1` is mandatory on a T4 β batch>1 OOMs. Gradient accumulation (4) | |
| recovers an effective batch of 4. | |
| ## Results so far (smoke test, 1000 samples, 10 epochs, T4) | |
| | Metric | Value | | |
| |---|---| | |
| | Train loss (final) | 4.6753 | | |
| | Eval loss (final) | 5.3534 | | |
| **The model has not learned reasoning.** Generated text reproduces the surface | |
| format of the training data ("Let me analyze this problem step by step") but the | |
| content is not grounded in the input. | |
| ## Open questions | |
| 1. **Is the conditioning doing anything?** `ablation.py` has not been run. Until | |
| it is, the contribution of Data A + Data B is unverified. This is the single | |
| most important experiment for this architecture. | |
| 2. **Tied embedding.** The output head reuses the encoder's frozen embedding | |
| matrix, saving ~132M params. But that matrix was never trained for | |
| generation, so the output geometry may be a poor fit. An untied head is worth | |
| trying. | |
| 3. **Data B source.** Currently the last hidden state of the 2-layer backbone. | |
| Intermediate layers, or the attention-head layers, may carry better signal. | |
| 4. **Scale.** 1000 samples / 10 epochs is a smoke test, not a training run. | |
| Eval loss 5.35 is far from usable. | |