French → English Transformer (trained from scratch)
An encoder-decoder Transformer trained from scratch (no pretrained weights) for fr→en translation, built for a generalization study: in-domain vs long sentences vs an unseen domain (literature).
Code, training and the full write-up: https://github.com/bad2212/transfromer-model. Training curves: https://wandb.ai/badalthakur2212-iisc/fr-en-transformer/runs/5ks4crpz.
Architecture
| Type | encoder-decoder, pre-LayerNorm (Xiong et al. 2020) |
| Layers | 6 encoder / 3 decoder (deep-encoder/shallow-decoder, Kasai et al. 2021) |
| Width | d_model 512, 8 heads, FFN 2048, ReLU, dropout 0.1 |
| Positions | rotary (RoPE, Su et al. 2021) in self-attention |
| Embeddings | one matrix shared by encoder input, decoder input and output layer (Press & Wolf 2017) |
| Parameters | 39.7M |
| Tokenizer | SentencePiece unigram, 16000 joint fr+en pieces, byte fallback; source-side subword sampling α=0.1 during training |
| Decoding | beam 5, GNMT length penalty α=1.8; inputs over 128 tokens are split at sentence boundaries |
Training data
Helsinki-NLP/opus-100 (en-fr), train split, fr→en. After normalisation (NFKC, typographic quotes to
ASCII), filtering (empty, source = target, length ratio outside [0.4, 2.5],
over 1000 chars), deduplication and removal of any pair whose source appears in the
dev/test sets: 923,274 of 1,000,000 pairs kept. opus_books was never used
for training or model selection. Label smoothing 0.1, AdamW, warmup 4000 + inverse-sqrt, ~24k tokens/update,
fp16 on one T4; final weights = average of the last checkpoints.
Results (provided dev set, official score.py, lowercased)
| slice | n | BLEU | chrF |
|---|---|---|---|
| seen | 60 | 29.78 | 46.37 |
| long | 30 | 34.66 | 59.98 |
| unseen_domain | 60 | 21.89 | 43.41 |
| all | 150 | 30.23 | 47.91 |
OVERALL (0.4·BLEU + 0.4·chrF + 0.2·chrF on unseen_domain) = 39.94
Larger eval-only sets (post-hoc, not used for selection)
| set | n | BLEU | chrF |
|---|---|---|---|
| opus100_test | 1000 | 31.56 | 52.02 |
| opus_books_sample | 1000 | 15.82 | 37.92 |
Dev slices are small (60/30/60 sentences), so per-slice numbers carry wide confidence intervals (see the report).
Usage
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("Badalt/fr-en-transformer-scratch")
sys.path.insert(0, path)
from fr_en_transformer import load_translator # needs: torch, sentencepiece, safetensors
tr = load_translator(path) # beam=5, alpha=1.8; override e.g. beam=1
print(tr.translate(["Ne t'inquiète pas !", "Il jeta son chapeau par terre."]))
Limitations
- Trained only on opus-100 (largely subtitles, web and official texts), so it does best on short, conversational or administrative sentences. Literary French (archaic vocabulary, long dependencies) is clearly weaker: see the unseen_domain slice.
- Names and rare words are split into subwords or bytes and are often mistranslated or transliterated.
- Single-reference automatic metrics with the challenge tokenisation; not comparable to sacreBLEU.
- fr→en only; outputs keep case but use ASCII quotes.
- Downloads last month
- 14