ukung's picture
Add source code (encoder_loader, model, data, train, ablation, generate) + NOTES
fd3090c verified
|
Raw History Blame Contribute Delete
1.87 kB
# Code — Semantic-Lite-2 Decoder
Source for the smoke-test model at
[`ukung/semantic-lite-2-decoder-smoke-test`](https://huggingface.co/ukung/semantic-lite-2-decoder-smoke-test).
## Files
| File | Purpose |
|---|---|
| `encoder_loader.py` | Loads the frozen Semantic-Lite-2 encoder + applies the two 4.57.1 compatibility patches. Extracts Data A / Data B. |
| `model.py` | `SemanticConditionedDecoder` (main model) and `DecoderNoConditioning` (ablation baseline). |
| `data.py` | Dataset loading, target formatting, tokenization, batching. |
| `train.py` | Training loop. Run this to reproduce the smoke test. |
| `ablation.py` | Conditioned vs unconditioned comparison. **Not yet run.** |
| `generate.py` | Inference from a prompt. |
| `NOTES.md` | Gotchas, current results, open questions. **Read this first.** |
| `requirements.txt` | Pinned deps (`transformers==4.57.1` is mandatory). |
## Quick start (Colab)
```python
!pip install -q -r requirements.txt
# then, with this folder on the path:
!python train.py
```
Or in a notebook:
```python
from encoder_loader import load_encoder, set_train_mode
from model import SemanticConditionedDecoder
from data import build_dataset
encoder, tokenizer = load_encoder(device="cuda")
model = SemanticConditionedDecoder(encoder, tokenizer).cuda()
```
## Architecture
```
input text
-> Semantic-Lite-2 (FROZEN)
Data A: (B, 256) -> proj_a -> prefix token
Data B: (B, L, 2048) -> proj_b -> cross-attention key/value
-> TransformerDecoderLayer (d_model=1024, nhead=16, FFN=1024) [TRAINABLE]
-> output_proj (1024 -> 2048) + tied embedding (131072 vocab)
```
Trainable parameters: **19,154,944**. The encoder is fully frozen.
## Status
Smoke test only. Train loss 4.6753 / eval loss 5.3534 after 10 epochs on 1000
samples. The model has **not** learned reasoning — see `NOTES.md`.