v5: actor-representation fix (frozen DINOv2 patch feats + frozen SmolLM2 text embedding into the actor), text-aware WM, BC-then-RL; CPU smoke passed all phases
v3: ent_coef 0.01 -> 0.05 to stop policy entropy collapse (v2 diagnosis: WM generalizes fine, val CE == train CE; actor entropy hit 0.0 by RL step 2666)
v3 backbone version: frozen DINOv2-small encoder + 36-token 6-dim FSQ (70 bytes/frame), SmolLM2-135M world model with real text instructions, KV-cache imagination
v2: anti-memorization (90/10 val split + WM early stopping), 2-head reward ensemble with pessimistic min in imagination, Lc=3 contexts, advantage normalization, 3-task demos