Neurograf-560M-Base
Neurograf-560M-Base is an approximately 560-million-parameter decoder-only language model pretrained from scratch on 131 billion Romanian-language tokens.
It belongs to Neurograf v1, an open research initiative focused on developing Romanian-first language models, tokenizers, and training resources.
This is a base language model, not an instruction-tuned assistant. It is intended for text continuation, language modeling research, and further fine-tuning.
Model Architecture
Neurograf uses a modified Nanochat transformer architecture with rotary positional embeddings, QK normalization, parameter-free RMSNorm, ReLU² feed-forward activations, and untied input/output embeddings.
| Parameter | Value |
|---|---|
| Architecture | Decoder-only Transformer |
| Parameters | Approximately 560M |
| Transformer layers | 12 |
| Hidden dimension | 1,536 |
| Attention heads | 12 |
| Key/value heads | 12 |
| Head dimension | 128 |
| Vocabulary size | 80,192 |
| Training sequence length | 2,048 |
| Positional encoding | RoPE |
| Attention | Multi-head causal attention |
| Activation | ReLU² |
| Normalization | RMSNorm, QK normalization |
| Output logits | Tanh softcap (15) |
| Training framework | Modified Nanochat |
| Checkpoint | Step 250,000 |
The architecture supports grouped-query attention, although this model uses standard multi-head attention with equal numbers of query and key/value heads.
Training Data
Neurograf-560M-Base was pretrained on approximately 131 billion tokens of Romanian-language text, using the Neurograf tokenizer developed specifically for the project.
The training corpus contains approximately 170 million Romanian text blocks collected and processed for language modeling.
All five Neurograf v1 base models were trained using the same overall token budget, enabling direct comparisons across model sizes.
Pretraining Results
Training Diagnostics
The following figure presents training loss, gradient norms, and validation bits per byte (BPB) for the five Neurograf v1 base models.
Across the model family, larger models generally achieve lower validation BPB, with continued improvement throughout pretraining.
The final 560M checkpoint achieves:
Validation BPB: 0.742747
The checkpoint was saved at training step 250,000, corresponding to approximately 131.1 billion training-token positions.
This checkpoint also achieves the lowest recorded validation BPB for its training run, as reported in the checkpoint metadata.
Fixed-Data Scaling Law
We examine how validation performance scales with model size when the training token budget is held fixed at approximately 131 billion tokens.
The empirical relationship is described by:
BPB(N) = 0.437 + 144 × N^(-0.305)
with a reported fit of R² = 0.99999 across the five model sizes.
This relationship describes the observed Neurograf v1 results under a fixed training-data budget. Extrapolations beyond the measured model sizes should be treated as exploratory rather than established performance predictions.
Model Files
The repository contains the original PyTorch checkpoint and tokenizer artifacts:
| File | Description |
|---|---|
model_250000.pt |
Final pretrained model weights |
meta_250000.json |
Architecture configuration and training metadata |
tokenizer.pkl |
Original tiktoken-based BPE tokenizer |
token_bytes.pt |
Token-byte mapping artifact |
The model is distributed in its original Nanochat-compatible checkpoint format. It has not yet been converted into a standard Hugging Face Transformers model.
Loading requires a compatible implementation of the modified Neurograf/Nanochat architecture. Inference examples and additional compatibility tools may be published separately.
Security note: The original .pt and .pkl formats use Python serialization mechanisms. Load these files only from trusted sources.
Intended Uses
- Romanian-language modeling research
- Continued pretraining
- Supervised fine-tuning
- Evaluation of Romanian language capabilities
- Model scaling and training dynamics research
- Development of specialized Romanian-language models
Limitations
This is a pretrained base model and has not been optimized for conversational instruction following.
Generated text may contain factual inaccuracies, biases, repetitions, or inappropriate content. The model has not been comprehensively evaluated for safety or reliability in high-stakes applications.
The training context length is 2,048 tokens. Longer-context behavior has not been systematically validated.
License
Released under the Apache License 2.0, subject to applicable third-party rights and attribution requirements.
Project
Neurograf — Developing monolingual Romanian language models from scratch.
Citation
If you use Neurograf in your research, please cite the Neurograf v1 technical report:
Neurograf v1: An End-to-End Romanian-First Language Model Family.
Full bibliographic information and a BibTeX citation can be added here.

