ukung's picture
Add source code (encoder_loader, model, data, train, ablation, generate) + NOTES
fd3090c verified
|
Raw History Blame Contribute Delete
3.08 kB

NOTES — gotchas and current state

Everything below was hit for real while building this. Read it before changing anything, especially the encoder loading.

Environment

transformers must be pinned to 4.57.1. Version 5.x breaks the Semantic-Lite-2 custom code in three separate ways (create_causal_mask signature, the rope validator, tied-weight handling). requirements.txt pins it.

Encoder loading (the tricky part)

ukung/semantic-lite-2 uses custom modelling code, so AutoModel.from_pretrained does not work. Two source patches are required, both applied at load time in encoder_loader.py:

  1. "inputs_embeds" → "input_embeds" — the keyword create_causal_mask expects in 4.57.1.
  2. Add "cache_position": cache_position to mask_kwargs — required in 4.57.1.

Also: rope_parameters=None is passed to the config to bypass the validator, then set manually afterwards. Without this the config constructor rejects the rope settings.

After patching, sys.modules entries starting with semantic_lite are purged so the patched file is what actually gets imported. Skipping that step silently loads a stale module.

Frozen encoder must stay in eval()

model.train() recurses into every submodule — including the frozen encoder, whose attention heads carry dropout=0.1. That injects noise into Data A and Data B on every step. Always use set_train_mode(model) from encoder_loader.py, never a bare .train().

Token alignment

tgt_ids   = [t0, t1, ..., tN, EOS]     <- no BOS
dec_input = [BOS, t0, t1, ..., tN]     <- shifted right
labels    = [t0, t1, ..., tN, EOS]

EOS lands in dec_input at a position whose label is PAD. That is correct (the token before the first PAD is EOS) and no gradient flows there.

Memory

Logits are (B, L, 131072). At L=768 that is ~100M floats per sample, so batch_size=1 is mandatory on a T4 — batch>1 OOMs. Gradient accumulation (4) recovers an effective batch of 4.

Results so far (smoke test, 1000 samples, 10 epochs, T4)

Metric Value
Train loss (final) 4.6753
Eval loss (final) 5.3534

The model has not learned reasoning. Generated text reproduces the surface format of the training data ("Let me analyze this problem step by step") but the content is not grounded in the input.

Open questions

  1. Is the conditioning doing anything? ablation.py has not been run. Until it is, the contribution of Data A + Data B is unverified. This is the single most important experiment for this architecture.
  2. Tied embedding. The output head reuses the encoder's frozen embedding matrix, saving ~132M params. But that matrix was never trained for generation, so the output geometry may be a poor fit. An untied head is worth trying.
  3. Data B source. Currently the last hidden state of the 2-layer backbone. Intermediate layers, or the attention-head layers, may carry better signal.
  4. Scale. 1000 samples / 10 epochs is a smoke test, not a training run. Eval loss 5.35 is far from usable.