Download code/NOTES.md from ukung/semantic-lite-2-decoder-smoke-test: direct link, hf CLI and curl.
- Browser
- Download file 3.08 kB
-
https://huggingface.co/ukung/semantic-lite-2-decoder-smoke-test/resolve/main/code/NOTES.md
- Command line
-
hf download hf://ukung/semantic-lite-2-decoder-smoke-test/code/NOTES.md
-
curl -L -o NOTES.md https://huggingface.co/ukung/semantic-lite-2-decoder-smoke-test/resolve/main/code/NOTES.md
NOTES — gotchas and current state
Everything below was hit for real while building this. Read it before changing anything, especially the encoder loading.
Environment
transformers must be pinned to 4.57.1. Version 5.x breaks the
Semantic-Lite-2 custom code in three separate ways (create_causal_mask
signature, the rope validator, tied-weight handling). requirements.txt pins it.
Encoder loading (the tricky part)
ukung/semantic-lite-2 uses custom modelling code, so AutoModel.from_pretrained
does not work. Two source patches are required, both applied at load time in
encoder_loader.py:
"inputs_embeds"→"input_embeds"— the keywordcreate_causal_maskexpects in 4.57.1.- Add
"cache_position": cache_positiontomask_kwargs— required in 4.57.1.
Also: rope_parameters=None is passed to the config to bypass the validator,
then set manually afterwards. Without this the config constructor rejects the
rope settings.
After patching, sys.modules entries starting with semantic_lite are purged
so the patched file is what actually gets imported. Skipping that step silently
loads a stale module.
Frozen encoder must stay in eval()
model.train() recurses into every submodule — including the frozen encoder,
whose attention heads carry dropout=0.1. That injects noise into Data A and
Data B on every step. Always use set_train_mode(model) from
encoder_loader.py, never a bare .train().
Token alignment
tgt_ids = [t0, t1, ..., tN, EOS] <- no BOS
dec_input = [BOS, t0, t1, ..., tN] <- shifted right
labels = [t0, t1, ..., tN, EOS]
EOS lands in dec_input at a position whose label is PAD. That is correct (the
token before the first PAD is EOS) and no gradient flows there.
Memory
Logits are (B, L, 131072). At L=768 that is ~100M floats per sample, so
batch_size=1 is mandatory on a T4 — batch>1 OOMs. Gradient accumulation (4)
recovers an effective batch of 4.
Results so far (smoke test, 1000 samples, 10 epochs, T4)
| Metric | Value |
|---|---|
| Train loss (final) | 4.6753 |
| Eval loss (final) | 5.3534 |
The model has not learned reasoning. Generated text reproduces the surface format of the training data ("Let me analyze this problem step by step") but the content is not grounded in the input.
Open questions
- Is the conditioning doing anything?
ablation.pyhas not been run. Until it is, the contribution of Data A + Data B is unverified. This is the single most important experiment for this architecture. - Tied embedding. The output head reuses the encoder's frozen embedding matrix, saving ~132M params. But that matrix was never trained for generation, so the output geometry may be a poor fit. An untied head is worth trying.
- Data B source. Currently the last hidden state of the 2-layer backbone. Intermediate layers, or the attention-head layers, may carry better signal.
- Scale. 1000 samples / 10 epochs is a smoke test, not a training run. Eval loss 5.35 is far from usable.