Koine-Greek-BERT (v1.0.1)

A domain-adapted BERT for Koine Greek, continued-pretrained from pranaydeeps/Ancient-Greek-BERT with a 50,000-token vocabulary extended for Koine word forms.

The model is accent-insensitive by design. It reads polytonic and monotonic input interchangeably — λόγος and λογος produce the identical token id — which makes it robust to differences of editorial accentuation between critical editions. It does not represent accentuation: diacritics are discarded during normalisation and cannot be recovered from the model's output.

Tokenizer correction, 29 September 2026 (v1.0.1). Model weights unchanged. In v1.0, 2,476 of the 15,000 added tokens become identical to another added token once accents are stripped. Transformers 4.x resolved each such collision to the lowest id, which is the tokenisation the model was trained and evaluated with. Transformers 5.x resolves the same collisions in an order that depends on Python's hash seed, so the same word can receive a different id in every process. Under transformers 5.17 the v1.0 files therefore give non-reproducible input ids and an evaluation perplexity of about 4,460 instead of about 10 (see Evaluation). v1.0.1 retires the 2,476 duplicate entries: each keeps its id, and its label is replaced by an unmatchable placeholder ([unused_v1.0.1_<id>]). The surviving token in every collision group is the lowest id, as in transformers 4.x. Verified: on 3,000 validation lines and on 77,700 words of New Testament text, v1.0.1 under transformers 5.17 and under 4.57.3, and v1.0 under 4.57.3, produce identical input ids, and v1.0.1 produces the same ids under different hash seeds. Users of v1.0 with transformers ≥ 5 should update the tokenizer files; results obtained with the v1.0 tokenizer under transformers 5.x should be recomputed.

Usage note, 30 September 2026 (no change to the model or tokenizer files). Because added tokens are also matched inside words (see Known limitations of the vocabulary), a sentence passed to the tokenizer as one string yields word_ids() in which the fragments of a split word count as separate words. Code that averages sub-word vectors per word and expects one vector per whitespace-separated word will then fail or silently misalign vectors and words. Pass pre-split words with is_split_into_words=True (see How to use). The input ids are identical in both modes; only the word mapping differs.

Documentation correction, 23 September 2026. Earlier versions of this card described the vocabulary as "50,000 (Polytonic)" and the model as adapted to "full polytonic accentuation". That description was inaccurate. The 15,000 added tokens were trained on polytonic text but registered with normalized: true, so their accent-stripped forms are what the tokenizer matches.

Model details

Base model pranaydeeps/Ancient-Greek-BERT
Architecture BERT, 12 layers, 768 hidden, 12 attention heads, 512 max positions
Parameters 124,492,880
Vocabulary 50,000 ids (35,000 base WordPiece + 15,000 added + 5 special); in v1.0.1, 2,476 duplicate added ids are retired (see below)
Objective Masked language modelling
Precision float32
Language Koine Greek (grc)

Tokenisation

The normaliser is BertNormalizer with lowercase=true; strip_accents is unset and therefore follows lowercase, so all diacritics are removed before matching. The added tokens carry normalized: true, so they too are matched in accent-stripped form.

What the vocabulary extension achieved

Compared with v0.1, the extension is a substantial improvement for Koine:

Input v0.1 v1.0 / v1.0.1
λόγος ['[UNK]'] ['λογος']
ἐν ἀρχῇ ἦν ὁ λόγος five [UNK] ['εν','αρχη','ην','ο','λογος']
ουν (οὖν) ['ο','##υ','##ν'] single token
μεν (μέν) ['με','##ν'] single token

v0.1 rejected accented input outright and fragmented the Greek discourse particles. v1.0 accepts both forms and gives οὖν, μέν, οὐ, τίς and ἦν whole-word tokens — which matters for any analysis that relies on function words and particles.

Known limitations of the vocabulary

The 15,000 added tokens comprise 11,016 whole-word strings and 3,984 strings beginning with ##. After accent-stripping:

  • 12,524 are distinct forms and 2,476 are duplicates of another added token (1,759 whole-word duplicates in 1,285 collision groups, and 717 ## duplicates). In v1.0.1 the duplicates are retired and the lowest id of each group is kept (see the correction note above).
  • 2,814 distinct forms coincide with an entry of the 35,000-token base vocabulary. For the 2,183 whole-word forms among them the base id is no longer produced.
  • 9,710 are genuinely new forms.
  • The 3,984 ## entries are never matched. Added tokens are matched as literal strings in the normalised text, where ## does not occur; the WordPiece continuation pieces of the base vocabulary are unaffected.
  • Added tokens are matched inside words. A word containing an added form is split at that form before WordPiece applies, e.g. κλητὸς → ['κλη', 'το', 'ς'], ἀφωρισμένος → ['αφωρισ', 'μεν', 'ος']. In the 77,700 words of the Greek New Testament letters tested, 20.4% are split in this way. The model was trained with this behaviour, so v1.0.1 keeps it. Changing it (for example by registering the added tokens as single_word) would alter the input ids and move the input away from the training distribution, so it is not offered as a correction.
  • Consequence for word-level work. When a whole sentence is tokenised as one string, each fragment produced by an in-word match is reported by word_ids() as a separate word. On 518 windows of 150 words each from the Pauline and other New Testament letters, string-mode tokenisation yields 209 word units per window on average and a mismatch in every window; v0.1 yields exactly 150. With is_split_into_words=True the word mapping is correct and the input ids are identical to string mode in all 518 windows (checked with tokenizers 0.23, 30 September 2026). The problem was reported for a word-alignment pipeline that tokenises whole verses (491 of 552 sentences affected).

Because the surviving id in each collision group is the lowest one, the human-readable label attached to a token need not be the most frequent spelling. For example, five spellings of καί collide — καὶ (35244), καί (35583), καῖ (37615), καϊ (40692), καἰ (44466) — and the id produced is 35244. (The previous version of this card gave 40692; that was the outcome of one transformers 5.x process and is not stable.) Decoded output should not be read as evidence about accentuation, and qualitative tables generated from convert_ids_to_tokens should be normalised before publication.

How to use

For sentence-level use (masked-language modelling, pooled sentence vectors) the tokenizer can be called on plain strings. For anything that maps sub-word vectors back to words, pass pre-split words:

from transformers import AutoTokenizer, AutoModel
tok = AutoTokenizer.from_pretrained("ABeZet/Koine-Greek-BERT", revision="v1.0.1")
model = AutoModel.from_pretrained("ABeZet/Koine-Greek-BERT", revision="v1.0.1")
words = "Παῦλος κλητὸς ἀπόστολος ἀφωρισμένος".split()
enc = tok(words, is_split_into_words=True, return_tensors="pt")
enc.word_ids()   # [None, 0, 1, 1, 1, 2, 3, 3, 3, None]: one index per input word

Training data

Trained on an assembled corpus of 177,608 lines, split 168,727 train / 8,881 validation. The corpus itself is not redistributed; it was built from the following openly licensed sources and can be reconstructed from them.

Source Contribution Licence
MorphGNT / SBLGNT Greek New Testament CC BY-SA 3.0
Byzantine Majority Text Greek New Testament, Byzantine tradition CC BY 4.0
Swete Septuagint Greek Old Testament CC BY-SA 4.0
First1KGreek Philo, Apostolic Fathers, Origen, Eusebius, Clement of Alexandria, Athanasius, Gregory of Nazianzus, Epiphanius CC BY-SA 4.0
Perseus Digital Library Josephus, Polybius, Epictetus, Appian, Diodorus Siculus, Strabo CC BY-SA 3.0
Gorman Dependency Trees Classical Attic and Ionic prose CC BY 4.0
PROIEL Treebank Greek New Testament, Herodotus, a Byzantine chronicle CC BY-NC-SA 4.0

Composition by content, computed from the corpus manifest:

Component Lines Share
Patristic (Origen, Eusebius, Epiphanius, Athanasius, Clement of Alexandria, Gregory of Nazianzus) 33,714 18.98%
Septuagint (Swete) 29,023 16.34%
New Testament (three text forms, see below) 27,400 15.43%
Secular Koine (Polybius, Diodorus Siculus, Strabo, Appian, Epictetus) 27,246 15.34%
Classical Attic and Ionic prose (Herodotus, Xenophon, Demosthenes, Thucydides, Plato, Aristotle, Attic orators) 25,849 14.55%
Jewish Koine (Philo, Josephus) 20,734 11.67%
Atticising prose of the Roman period (Plutarch, Dionysius of Halicarnassus and others) 9,358 5.27%
Apostolic Fathers (1 Clement, Hermas, Ignatius, Barnabas, Didache, Polycarp) 3,259 1.83%
Byzantine Greek (a chronicle internally dated to ca. 1400 CE) 1,008 0.57%
Latin (Cicero, Ad Atticum, ingested in error from the multilingual PROIEL corpus) 17 0.01%

Two points deserve emphasis for anyone considering this model:

The New Testament is 15.43% of the training data, present in three independent text forms — PROIEL (6.48%), the Byzantine majority text (4.49%) and SBLGNT/MorphGNT (4.46%). The model has therefore seen the New Testament repeatedly and in variant readings. It is not suitable, without explicit caveat, for authorship attribution, interpolation detection or any other evaluation carried out on New Testament text, because such an evaluation cannot separate stylistic signal from memorisation of the training data.

Roughly a fifth of the corpus is not Koine. Classical Attic and Ionic prose accounts for 14.55%, and atticising Roman-period prose for a further 5.27%. This was a deliberate choice when the corpus was assembled, intended to broaden syntactic coverage; it is recorded here so that the name of the model is not mistaken for a statement about the composition of its training data.

Training procedure

Two-phase domain-adaptive pre-training (DAPT), carried out with transformers 4.57:

  • Phase 1 (embedding warm-up): base weights frozen, 4 epochs, learning rate 5e-4. New token embeddings were initialised by averaging the embeddings of their sub-word components in the base model.
  • Phase 2 (full tuning): all weights unfrozen, 4 epochs, learning rate 5e-5 with 200 warm-up steps, float32, effective batch size 64 (2 × T4, batch 4, gradient accumulation 8). float16 was attempted and abandoned after NaN gradient overflow on T4.

Evaluation

Metric Phase 1 Phase 2 (v1.0)
Evaluation loss 2.5946 2.3423
Perplexity 13.39 10.41

Measured on val_v1.txt (8,881 lines), CPU, single device, transformers 4.57.

Check of the tokenizer correction (29 September 2026). Same weights, transformers 5.17.

Tokenizer Validation lines Evaluation loss Perplexity
v1.0 files (hash seed 1) 400 8.40 ≈ 4,460
v1.0.1 files 400 2.30 10.00
v1.0.1 files 8,881 (full set) 2.3448 10.43

The full-set figure follows the procedure of the original evaluation (maximum length 512, batches of 16, 15% random masking, mean of batch losses; masking seeded by batch number) and reproduces the v1.0 result of May 2026 (2.3423, 10.41) within the variation of random masking.

Caveat on the split. The validation set was produced by shuffling lines and taking 5%, so it draws on the same works as the training set. The reported perplexity therefore measures fit to the training distribution, not generalisation to unseen works. A held-out-by-work split would give a higher and more informative figure.

Zero-shot cloze probes (v1.0.1, transformers 5.17): ἐν ἀρχῇ ἦν ὁ λόγος , καὶ ὁ λόγος ἦν πρὸς τὸν θεόν , καὶ θεὸς ἦν ὁ [MASK] . → λόγος at rank 1; ἀγαπήσεις τὸν [MASK] σου ὡς σεαυτόν . → πλησίον at rank 1. Note that the accents shown in such output are stored labels, not predictions (see Known limitations of the vocabulary above).

Intended use

Suitable for masked-language-modelling over Koine and related Ancient Greek prose, for feature extraction, and as a starting point for fine-tuning on downstream tasks over non-New-Testament Greek.

Not suitable, without additional controls, for: evaluation on New Testament text (training-data overlap); any task requiring accentual or prosodic information (discarded at normalisation); tasks requiring faithful orthographic reconstruction from token labels.

Licence

Released under CC BY-NC-SA 4.0: attribution, non-commercial use, share-alike.

The non-commercial condition is inherited from the training data. The PROIEL Treebank, which contributed 18,190 lines (10.24% of the corpus), is licensed CC BY-NC-SA 4.0 and carries a non-commercial clause. This model is accordingly released for research use only.

Roadmap

A successor model focused on authorship-sensitive representation of Ancient Greek is under development as part of the NEURAL SCRIBE project (NCN SONATINA 10). Unlike the present model, it will exclude the New Testament from its training data, so that New Testament text remains available for evaluation, and will be built without non-commercially licensed sources. No release date is fixed.

Availability of the training corpus

The assembled corpus is not published. Its sources carry several different licences, including one non-commercial and several share-alike, and a redistribution has not been cleared. The source list and composition table above are given in full so that the training data can be assessed, and reconstructed from the original repositories, without access to the assembled file.

Version history

  • Model card, 30 September 2026. Usage note and How to use section on word-level tokenisation (is_split_into_words=True). No change to the model or tokenizer files.
  • v1.0.1 (current, 29 September 2026). Tokenizer files corrected so that collisions between added tokens are resolved deterministically to the lowest id, reproducing the tokenisation used in training under every transformers version. 2,476 duplicate entries retired (ids kept, labels replaced by placeholders). Model weights and configuration unchanged.
  • v1.0. Base model pranaydeeps/Ancient-Greek-BERT; vocabulary extended to 50,000 with smart-initialised embeddings; two-phase DAPT. Documentation corrected 23 September 2026.
  • v0.1 (deprecated). Proof of concept on nlpaueb/bert-base-greek-uncased-v1 (Modern Greek BERT), 35,000-token vocabulary, polytonic input not accepted.

Citation

@misc{zieminska_koine_greek_bert,
  author       = {Ziemińska, Agnieszka},
  title        = {Koine-Greek-BERT},
  year         = {2026},
  publisher    = {Hugging Face},
  doi          = {10.57967/hf/10565},
  howpublished = {\url{https://huggingface.co/ABeZet/Koine-Greek-BERT}}
}
Downloads last month
62
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support