# NOTES — gotchas and current state Everything below was hit for real while building this. Read it before changing anything, especially the encoder loading. ## Environment **`transformers` must be pinned to 4.57.1.** Version 5.x breaks the Semantic-Lite-2 custom code in three separate ways (`create_causal_mask` signature, the rope validator, tied-weight handling). `requirements.txt` pins it. ## Encoder loading (the tricky part) `ukung/semantic-lite-2` uses custom modelling code, so `AutoModel.from_pretrained` does not work. Two source patches are required, both applied at load time in `encoder_loader.py`: 1. `"inputs_embeds"` → `"input_embeds"` — the keyword `create_causal_mask` expects in 4.57.1. 2. Add `"cache_position": cache_position` to `mask_kwargs` — required in 4.57.1. Also: `rope_parameters=None` is passed to the config to bypass the validator, then set manually afterwards. Without this the config constructor rejects the rope settings. After patching, `sys.modules` entries starting with `semantic_lite` are purged so the patched file is what actually gets imported. Skipping that step silently loads a stale module. ## Frozen encoder must stay in eval() `model.train()` recurses into every submodule — including the frozen encoder, whose attention heads carry `dropout=0.1`. That injects noise into Data A and Data B on every step. Always use `set_train_mode(model)` from `encoder_loader.py`, never a bare `.train()`. ## Token alignment ``` tgt_ids = [t0, t1, ..., tN, EOS] <- no BOS dec_input = [BOS, t0, t1, ..., tN] <- shifted right labels = [t0, t1, ..., tN, EOS] ``` EOS lands in `dec_input` at a position whose label is PAD. That is correct (the token before the first PAD is EOS) and no gradient flows there. ## Memory Logits are `(B, L, 131072)`. At L=768 that is ~100M floats per sample, so `batch_size=1` is mandatory on a T4 — batch>1 OOMs. Gradient accumulation (4) recovers an effective batch of 4. ## Results so far (smoke test, 1000 samples, 10 epochs, T4) | Metric | Value | |---|---| | Train loss (final) | 4.6753 | | Eval loss (final) | 5.3534 | **The model has not learned reasoning.** Generated text reproduces the surface format of the training data ("Let me analyze this problem step by step") but the content is not grounded in the input. ## Open questions 1. **Is the conditioning doing anything?** `ablation.py` has not been run. Until it is, the contribution of Data A + Data B is unverified. This is the single most important experiment for this architecture. 2. **Tied embedding.** The output head reuses the encoder's frozen embedding matrix, saving ~132M params. But that matrix was never trained for generation, so the output geometry may be a poor fit. An untied head is worth trying. 3. **Data B source.** Currently the last hidden state of the 2-layer backbone. Intermediate layers, or the attention-head layers, may carry better signal. 4. **Scale.** 1000 samples / 10 epochs is a smoke test, not a training run. Eval loss 5.35 is far from usable.