usrnotfound101/anlp-a2-moe-dense

Decoder-only Transformer trained from scratch to translate Vietnamese and Japanese -> English (ANLP Assignment 2, Part 1). FFN variant: dense - Standard 2-layer MLP (d -> 4d -> d).

value
total parameters 35,396,096
active parameters / token 35,396,096
FFN params (total / active) 12,582,912 / 12,582,912
layers / d_model / heads 6 / 512 / 8
training tokens (non-pad) 36,824,882
test perplexity (all / vi / ja) 7.893 / 6.574 / 9.476
test BLEU (all / vi / ja) 29.89 / 34.09 / 25.41
test chrF (all) 51.58

Prompt format: <vi> source <en> or <ja> source <en>; the model continues with the English translation and </s>. Tokenizer: joint SentencePiece unigram (mt_spm.model, byte fallback, no romanisation).

Files: model.safetensors, config.json, results.json, train_log.json, expert_usage.json, plots (*.png) and the model code (common.py). Load with:

from common import load_model_dir
model = load_model_dir("usrnotfound101/anlp-a2-moe-dense")
Downloads last month
8
Safetensors
Model size
35.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train usrnotfound101/anlp-a2-moe-dense