usrnotfound101/anlp-a2-opt-adamw

Dense decoder-only Transformer (27.40M params) pretrained from scratch on HAP-E for 1.0 x the train split with the adamw optimizer (AdamW (baseline)), implemented from scratch (see optimizers.py).

  • hyper-parameters: {"lr": 0.001, "betas": [0.9, 0.95], "eps": 1e-08, "weight_decay": 0.1, "grad_clip": 1.0}
  • final validation loss: 3.8301 (ppl 46.07)
  • final test BLEU (multi-reference continuation): 4.34
  • intermediate checkpoints every 0.1 x dataset tokens under checkpoints/
  • full training / evaluation log: train_log.json
Downloads last month
105
Safetensors
Model size
27.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train usrnotfound101/anlp-a2-opt-adamw