Neurograf-6B-Base
Neurograf-6B-Base is a 5.9-billion-parameter decoder-only language model pretrained from scratch on approximately 131 billion Romanian-language tokens.
It is part of Neurograf v1, an open research initiative focused on developing Romanian-first language models, tokenizers, and training resources.
The model is a base language model, not an instruction-tuned assistant. It is designed for text continuation, language modeling research, and further fine-tuning.
Model Architecture
Neurograf is built upon a modified version of NanoChat, the open-source language model training framework developed by Andrew Karpathy.
The Neurograf architecture retains NanoChat's lightweight, research-oriented design while adapting it for the development of Romanian-first language models. Its main architectural components include rotary positional embeddings (RoPE), QK normalization, parameter-free RMSNorm, ReLU² feed-forward activations, and untied input/output embeddings.
Reference: Karpathy, A. NanoChat. GitHub repository: https://github.com/karpathy/nanochat
| Parameter | Value |
|---|---|
| Architecture | Decoder-only Transformer |
| Parameters | Approximately 5.9B |
| Transformer layers | 30 |
| Hidden dimension | 3,840 |
| Attention heads | 30 |
| Key/value heads | 30 |
| Head dimension | 128 |
| Vocabulary size | 80,192 |
| Training sequence length | 2,048 |
| Positional encoding | RoPE |
| Attention | Multi-head causal attention |
| Activation | ReLU² |
| Normalization | RMSNorm, QK normalization |
| Output logits | Tanh softcap (15) |
| Training framework | Modified Nanochat |
| Training precision | Mixed precision |
| Checkpoint | Step 62,500 |
The architecture supports grouped-query attention, although this model uses equal numbers of query and key/value heads.
Training Data
Neurograf-6B-Base was pretrained on approximately 131 billion tokens of Romanian-language text, using the Neurograf tokenizer developed specifically for the project.
The training corpus contains approximately 170 million text blocks collected and processed for Romanian-language modeling.
All five Neurograf v1 base models were trained using the same overall token budget, enabling direct comparisons across model sizes.
Pretraining Results
Training Diagnostics
The following figure presents training loss, gradient norms, and validation bits per byte (BPB) for the five Neurograf v1 base models.
Across the model family, larger models generally achieve lower validation BPB, with continued improvement throughout pretraining.
The final 6B checkpoint achieves:
Validation BPB: 0.587616
Fixed-Data Scaling Law
We also examine how validation performance scales with model size when the training token budget is held fixed at approximately 131 billion tokens.
The empirical relationship is described by:
BPB(N) = 0.437 + 144 × N^(-0.305)
with a reported fit of R² = 0.99999 across the five model sizes.
This relationship describes the observed Neurograf v1 results under a fixed training-data budget. Extrapolations beyond the measured model sizes should be treated as exploratory rather than established performance predictions.
Model Files
The repository contains the original PyTorch checkpoint and tokenizer artifacts:
| File | Description |
|---|---|
model_062500.pt |
Final pretrained model weights |
meta_062500.json |
Architecture configuration and training metadata |
tokenizer.pkl |
Original tiktoken-based BPE tokenizer |
token_bytes.pt |
Token-byte mapping artifact |
The model is distributed in its original Nanochat-compatible checkpoint format. It has not yet been converted into a standard Hugging Face Transformers model.
Loading requires a compatible implementation of the modified Neurograf/Nanochat architecture. Inference examples and additional compatibility tools may be published separately.
Security note: The original .pt and .pkl formats use Python serialization mechanisms. Load these files only from trusted sources.
Intended Uses
- Romanian-language modeling research
- Continued pretraining
- Supervised fine-tuning
- Evaluation of Romanian language capabilities
- Model scaling and training dynamics research
- Development of specialized Romanian-language models
Limitations
This is a pretrained base model and has not been optimized for conversational instruction following.
Generated text may contain factual inaccuracies, biases, repetitions, or inappropriate content. The model has not been comprehensively evaluated for safety or reliability in high-stakes applications.
The training context length is 2,048 tokens. Longer-context behavior has not been systematically validated.
License
Released under the Apache License 2.0, subject to applicable third-party rights and attribution requirements.
Project
Neurograf — Developing monolingual Romanian language models from scratch.
Citation
If you use Neurograf in your research, please cite the Neurograf v1 technical report:
Neurograf v1: An End-to-End Romanian-First Language Model Family.
Full bibliographic information and a BibTeX citation can be added here.

