You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Apertus 8B Greek CPT — full checkpoint trajectory

Continued pretraining on 76.685B active tokens. Base models, not instruction tuned.

Training

79% Modern Greek, 20% foreign replay, 1% Old Greek. Extended vocabulary: 148,992 tokens. 18 training checkpoints and two checkpoint averages. main: terminal checkpoint.

Recommended checkpoint

base-avg30B-50B: GreekMMLU 56.78% (best single checkpoint: 56.81%). Apertus pretraining-suite macro: 64.58% (terminal: 62.95%).

Repositories and checkpoints

Resource Links
Original Apertus Repository, Revision
Instruction models Repository, Stage 1, Greek maths, Conversation
Greek base base-avg30B-50B, 17-step18284-tokens77B, base-avg30B-50B, CHECKPOINTS.md, checkpoint-index.json
Pre-training data Repository, Revision
SFT data Stage 1, Repository
Post-training data Greek maths, Conversation, Repository
Tokenizer Revision, Repository
Benchmark audit Repository

Acknowledgements

This work was implemented thanks to a grant by Swiss AI for compute on CSCS.

Collection: Greek Apertus 8B.

Downloads last month
8
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for glossAPI/apertus-8b-greek-cpt

Finetuned
(19)
this model
Finetunes
1 model

Dataset used to train glossAPI/apertus-8b-greek-cpt

Collection including glossAPI/apertus-8b-greek-cpt

Evaluation results

  • accuracy on GreekMMLU (decontaminated, n=16,159)
    self-reported
    54.850
  • accuracy on ASEP MCQA (strict contamination-filtered)
    self-reported
    55.080
  • accuracy on DemosQA (strict contamination-filtered)
    self-reported
    46.580
  • accuracy on GPCR (strict contamination-filtered)
    self-reported
    62.890
  • accuracy on Medical MCQA (strict contamination-filtered)
    self-reported
    38.420
  • accuracy on OYXOY metaphor (strict contamination-filtered)
    self-reported
    33.890
  • accuracy on OYXOY NLI (strict contamination-filtered)
    self-reported
    38.730
  • accuracy on OYXOY WiC (strict contamination-filtered)
    self-reported
    33.640