ukung's picture
Add source code (encoder_loader, model, data, train, ablation, generate) + NOTES
fd3090c verified
|
Raw History Blame Contribute Delete
1.87 kB

Code — Semantic-Lite-2 Decoder

Source for the smoke-test model at ukung/semantic-lite-2-decoder-smoke-test.

Files

File Purpose
encoder_loader.py Loads the frozen Semantic-Lite-2 encoder + applies the two 4.57.1 compatibility patches. Extracts Data A / Data B.
model.py SemanticConditionedDecoder (main model) and DecoderNoConditioning (ablation baseline).
data.py Dataset loading, target formatting, tokenization, batching.
train.py Training loop. Run this to reproduce the smoke test.
ablation.py Conditioned vs unconditioned comparison. Not yet run.
generate.py Inference from a prompt.
NOTES.md Gotchas, current results, open questions. Read this first.
requirements.txt Pinned deps (transformers==4.57.1 is mandatory).

Quick start (Colab)

!pip install -q -r requirements.txt
# then, with this folder on the path:
!python train.py

Or in a notebook:

from encoder_loader import load_encoder, set_train_mode
from model import SemanticConditionedDecoder
from data import build_dataset

encoder, tokenizer = load_encoder(device="cuda")
model = SemanticConditionedDecoder(encoder, tokenizer).cuda()

Architecture

input text
    -> Semantic-Lite-2 (FROZEN)
         Data A: (B, 256)      -> proj_a -> prefix token
         Data B: (B, L, 2048)  -> proj_b -> cross-attention key/value
    -> TransformerDecoderLayer (d_model=1024, nhead=16, FFN=1024)  [TRAINABLE]
    -> output_proj (1024 -> 2048) + tied embedding (131072 vocab)

Trainable parameters: 19,154,944. The encoder is fully frozen.

Status

Smoke test only. Train loss 4.6753 / eval loss 5.3534 after 10 epochs on 1000 samples. The model has not learned reasoning — see NOTES.md.