File size: 3,076 Bytes
fd3090c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
# NOTES — gotchas and current state

Everything below was hit for real while building this. Read it before changing
anything, especially the encoder loading.

## Environment

**`transformers` must be pinned to 4.57.1.** Version 5.x breaks the
Semantic-Lite-2 custom code in three separate ways (`create_causal_mask`
signature, the rope validator, tied-weight handling). `requirements.txt` pins it.

## Encoder loading (the tricky part)

`ukung/semantic-lite-2` uses custom modelling code, so `AutoModel.from_pretrained`
does not work. Two source patches are required, both applied at load time in
`encoder_loader.py`:

1. `"inputs_embeds"` → `"input_embeds"` — the keyword `create_causal_mask`
   expects in 4.57.1.
2. Add `"cache_position": cache_position` to `mask_kwargs` — required in 4.57.1.

Also: `rope_parameters=None` is passed to the config to bypass the validator,
then set manually afterwards. Without this the config constructor rejects the
rope settings.

After patching, `sys.modules` entries starting with `semantic_lite` are purged
so the patched file is what actually gets imported. Skipping that step silently
loads a stale module.

## Frozen encoder must stay in eval()

`model.train()` recurses into every submodule — including the frozen encoder,
whose attention heads carry `dropout=0.1`. That injects noise into Data A and
Data B on every step. Always use `set_train_mode(model)` from
`encoder_loader.py`, never a bare `.train()`.

## Token alignment

```
tgt_ids   = [t0, t1, ..., tN, EOS]     <- no BOS
dec_input = [BOS, t0, t1, ..., tN]     <- shifted right
labels    = [t0, t1, ..., tN, EOS]
```

EOS lands in `dec_input` at a position whose label is PAD. That is correct (the
token before the first PAD is EOS) and no gradient flows there.

## Memory

Logits are `(B, L, 131072)`. At L=768 that is ~100M floats per sample, so
`batch_size=1` is mandatory on a T4 — batch>1 OOMs. Gradient accumulation (4)
recovers an effective batch of 4.

## Results so far (smoke test, 1000 samples, 10 epochs, T4)

| Metric | Value |
|---|---|
| Train loss (final) | 4.6753 |
| Eval loss (final) | 5.3534 |

**The model has not learned reasoning.** Generated text reproduces the surface
format of the training data ("Let me analyze this problem step by step") but the
content is not grounded in the input.

## Open questions

1. **Is the conditioning doing anything?** `ablation.py` has not been run. Until
   it is, the contribution of Data A + Data B is unverified. This is the single
   most important experiment for this architecture.
2. **Tied embedding.** The output head reuses the encoder's frozen embedding
   matrix, saving ~132M params. But that matrix was never trained for
   generation, so the output geometry may be a poor fit. An untied head is worth
   trying.
3. **Data B source.** Currently the last hidden state of the 2-layer backbone.
   Intermediate layers, or the attention-head layers, may carry better signal.
4. **Scale.** 1000 samples / 10 epochs is a smoke test, not a training run.
   Eval loss 5.35 is far from usable.