French → English Transformer (trained from scratch)

An encoder-decoder Transformer trained from scratch (no pretrained weights) for fr→en translation, built for a generalization study: in-domain vs long sentences vs an unseen domain (literature).

Code, training and the full write-up: https://github.com/bad2212/transfromer-model. Training curves: https://wandb.ai/badalthakur2212-iisc/fr-en-transformer/runs/5ks4crpz.

Architecture

Type encoder-decoder, pre-LayerNorm (Xiong et al. 2020)
Layers 6 encoder / 3 decoder (deep-encoder/shallow-decoder, Kasai et al. 2021)
Width d_model 512, 8 heads, FFN 2048, ReLU, dropout 0.1
Positions rotary (RoPE, Su et al. 2021) in self-attention
Embeddings one matrix shared by encoder input, decoder input and output layer (Press & Wolf 2017)
Parameters 39.7M
Tokenizer SentencePiece unigram, 16000 joint fr+en pieces, byte fallback; source-side subword sampling α=0.1 during training
Decoding beam 5, GNMT length penalty α=1.8; inputs over 128 tokens are split at sentence boundaries

Training data

Helsinki-NLP/opus-100 (en-fr), train split, fr→en. After normalisation (NFKC, typographic quotes to ASCII), filtering (empty, source = target, length ratio outside [0.4, 2.5], over 1000 chars), deduplication and removal of any pair whose source appears in the dev/test sets: 923,274 of 1,000,000 pairs kept. opus_books was never used for training or model selection. Label smoothing 0.1, AdamW, warmup 4000 + inverse-sqrt, ~24k tokens/update, fp16 on one T4; final weights = average of the last checkpoints.

Results (provided dev set, official score.py, lowercased)

slice n BLEU chrF
seen 60 29.78 46.37
long 30 34.66 59.98
unseen_domain 60 21.89 43.41
all 150 30.23 47.91

OVERALL (0.4·BLEU + 0.4·chrF + 0.2·chrF on unseen_domain) = 39.94

Larger eval-only sets (post-hoc, not used for selection)

set n BLEU chrF
opus100_test 1000 31.56 52.02
opus_books_sample 1000 15.82 37.92

Dev slices are small (60/30/60 sentences), so per-slice numbers carry wide confidence intervals (see the report).

Usage

from huggingface_hub import snapshot_download
import sys
path = snapshot_download("Badalt/fr-en-transformer-scratch")
sys.path.insert(0, path)
from fr_en_transformer import load_translator   # needs: torch, sentencepiece, safetensors
tr = load_translator(path)                     # beam=5, alpha=1.8; override e.g. beam=1
print(tr.translate(["Ne t'inquiète pas !", "Il jeta son chapeau par terre."]))

Limitations

  • Trained only on opus-100 (largely subtitles, web and official texts), so it does best on short, conversational or administrative sentences. Literary French (archaic vocabulary, long dependencies) is clearly weaker: see the unseen_domain slice.
  • Names and rare words are split into subwords or bytes and are often mistranslated or transliterated.
  • Single-reference automatic metrics with the challenge tokenisation; not comparable to sacreBLEU.
  • fr→en only; outputs keep case but use ASCII quotes.
Downloads last month
14
Safetensors
Model size
39.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Badalt/fr-en-transformer-scratch