belumind/en-vi-ja-curated-500k-triplets
Viewer • Updated • 496k • 1.5k • 1
Advanced NLP (Monsoon 2026), Assignment 2.
Part 1, FFN variant 4. 8.8M total / 7.2M active params, trained on 120000979 tokens. Best checkpoint by validation loss; test loss 2.049, ppl 7.763, acc 0.626, bleu 22.811.
| item | value |
|---|---|
| parameters | 8.8M |
| training step | 89077 |
| wandb | https://wandb.ai/winterdew-iiit-hyderabad/anlp-a2-part1/runs/m5zatkbb |
| setting | value |
|---|---|
| FFN variant | 4 — 4 experts, 2 active per token (1 shared + 3 routed), total-param matched |
| params (total / active) | 8.8M / 7.2M |
| FFN hidden width | 256 per expert |
| MoE aux weight | 0.01 |
| d_model / n_layers / n_heads | 256 / 6 / 8 |
| context length | 256 |
| positional encoding | RoPE (base 10000) |
| tokenizer | sentencepiece unigram, vocab 16000, character coverage 0.9995, byte fallback |
| optimizer | AdamW(betas=(0.9, 0.95), eps=1e-8) |
| learning rate | 0.0003 |
| lr schedule | linear warmup 1740 steps, cosine to 0.1x |
| weight decay | 0.1 (matrices only) |
| gradient clip | 1.0 |
| batch size | 32 x 1 accum |
| dropout | 0.1 |
| precision | fp16 |
| token budget | 120,000,000 (120,000,979 seen) |
| seed | 26 |
{
"vocab_size": 16000,
"n_ctx": 256,
"rope_base": 10000.0,
"d_model": 256,
"n_heads": 8,
"n_layers": 6,
"d_ff": 1024,
"dropout": 0.1,
"tie_embeddings": true,
"use_moe": true,
"n_experts": 4,
"n_active": 2,
"n_shared": 1,
"match_active": false,
"moe_aux_weight": 0.01
}