belumind/en-vi-ja-curated-500k-triplets
Viewer • Updated • 496k • 1.04k • 1
Decoder-only Transformer trained from scratch to translate Vietnamese and Japanese -> English
(ANLP Assignment 2, Part 1). FFN variant: dense - Standard 2-layer MLP (d -> 4d -> d).
| value | |
|---|---|
| total parameters | 35,396,096 |
| active parameters / token | 35,396,096 |
| FFN params (total / active) | 12,582,912 / 12,582,912 |
| layers / d_model / heads | 6 / 512 / 8 |
| training tokens (non-pad) | 36,824,882 |
| test perplexity (all / vi / ja) | 7.893 / 6.574 / 9.476 |
| test BLEU (all / vi / ja) | 29.89 / 34.09 / 25.41 |
| test chrF (all) | 51.58 |
Prompt format: <vi> source <en> or <ja> source <en>; the model continues with the English
translation and </s>. Tokenizer: joint SentencePiece unigram (mt_spm.model, byte fallback, no romanisation).
Files: model.safetensors, config.json, results.json, train_log.json, expert_usage.json,
plots (*.png) and the model code (common.py). Load with:
from common import load_model_dir
model = load_model_dir("usrnotfound101/anlp-a2-moe-dense")