Text Generation
Transformers
Safetensors
Greek
apertus
greek
glossapi
continued-pretraining
Eval Results (legacy)
Instructions to use glossAPI/apertus-8b-greek-cpt with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use glossAPI/apertus-8b-greek-cpt with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="glossAPI/apertus-8b-greek-cpt")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("glossAPI/apertus-8b-greek-cpt") model = AutoModelForCausalLM.from_pretrained("glossAPI/apertus-8b-greek-cpt", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use glossAPI/apertus-8b-greek-cpt with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "glossAPI/apertus-8b-greek-cpt" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "glossAPI/apertus-8b-greek-cpt", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/glossAPI/apertus-8b-greek-cpt
- SGLang
How to use glossAPI/apertus-8b-greek-cpt with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "glossAPI/apertus-8b-greek-cpt" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "glossAPI/apertus-8b-greek-cpt", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "glossAPI/apertus-8b-greek-cpt" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "glossAPI/apertus-8b-greek-cpt", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use glossAPI/apertus-8b-greek-cpt with Docker Model Runner:
docker model run hf.co/glossAPI/apertus-8b-greek-cpt
Apertus 8B Greek CPT — full checkpoint trajectory
Continued pretraining on 76.685B active tokens. Base models, not instruction tuned.
Training
79% Modern Greek, 20% foreign replay, 1% Old Greek. Extended vocabulary: 148,992 tokens. 18 training checkpoints and two checkpoint averages. main: terminal checkpoint.
Recommended checkpoint
base-avg30B-50B: GreekMMLU 56.78% (best single checkpoint: 56.81%). Apertus pretraining-suite macro: 64.58% (terminal: 62.95%).
Repositories and checkpoints
| Resource | Links |
|---|---|
| Original Apertus | Repository, Revision |
| Instruction models | Repository, Stage 1, Greek maths, Conversation |
| Greek base | base-avg30B-50B, 17-step18284-tokens77B, base-avg30B-50B, CHECKPOINTS.md, checkpoint-index.json |
| Pre-training data | Repository, Revision |
| SFT data | Stage 1, Repository |
| Post-training data | Greek maths, Conversation, Repository |
| Tokenizer | Revision, Repository |
| Benchmark audit | Repository |
Acknowledgements
This work was implemented thanks to a grant by Swiss AI for compute on CSCS.
Collection: Greek Apertus 8B.
- Downloads last month
- 8
Model tree for glossAPI/apertus-8b-greek-cpt
Dataset used to train glossAPI/apertus-8b-greek-cpt
Viewer • Updated • 49.6M • 445
Collection including glossAPI/apertus-8b-greek-cpt
Evaluation results
- accuracy on GreekMMLU (decontaminated, n=16,159)self-reported54.850
- accuracy on ASEP MCQA (strict contamination-filtered)self-reported55.080
- accuracy on DemosQA (strict contamination-filtered)self-reported46.580
- accuracy on GPCR (strict contamination-filtered)self-reported62.890
- accuracy on Medical MCQA (strict contamination-filtered)self-reported38.420
- accuracy on OYXOY metaphor (strict contamination-filtered)self-reported33.890
- accuracy on OYXOY NLI (strict contamination-filtered)self-reported38.730
- accuracy on OYXOY WiC (strict contamination-filtered)self-reported33.640