Neurograf
AI & ML interests
Developing monolingual Romanian models.
Recent Activity
Neurograf v1
Developing Romanian language models from scratch.
Neurograf is an open research initiative focused on training Romanian-first language models, developing language resources, and making AI research accessible to everyone.
Neurograf v1 Model Family
Neurograf v1 consists of five decoder-only Transformer models, ranging from 160 million to approximately 6 billion parameters.
All five models were trained from scratch on approximately 131 billion Romanian-language tokens, using a shared Romanian-first tokenizer and a modified Nanochat architecture.
| Model | Parameters | Layers | Validation BPB |
|---|---|---|---|
| Neurograf-160M-Base | 160M | 6 | 0.886683 |
| Neurograf-560M-Base | 560M | 12 | 0.742747 |
| Neurograf-1.5B-Base | 1.5B | 18 | 0.665701 |
| Neurograf-3.2B-Base | 3.2B | 24 | 0.619530 |
| Neurograf-6B-Base | 6B | 30 | 0.587616 |
All models use a context length of 2,048 tokens and a shared vocabulary of 80,192 tokens.
These are pretrained base models intended for research, continued pretraining, and downstream fine-tuning.
Pretraining Results
The Neurograf v1 models provide a controlled comparison of language-model scaling across five architectures trained with approximately the same token budget.
Training Diagnostics
The training results show consistent improvements in validation bits per byte (BPB) as model size increases.
Fixed-Data Scaling
A power-law fit to the five measured model sizes achieves R² = 0.99999, illustrating the relationship between model capacity and validation performance under a fixed Romanian-language training budget.
Extrapolations beyond the evaluated model sizes remain exploratory.
Architecture and Tokenization
The Neurograf v1 models share a modified Nanochat Transformer architecture featuring:
- Rotary positional embeddings (RoPE)
- QK normalization
- RMSNorm
- ReLU² feed-forward activations
- Untied input and output embeddings
- A custom Romanian-first BPE tokenizer
The models are currently distributed as original PyTorch checkpoints with their corresponding architecture metadata and tokenizer files.
Datasets and Language Resources
In addition to pretrained models, Neurograf develops datasets and resources for Romanian-language training and adaptation.
- EN-RO Translation SFT 300K — English-to-Romanian translation training pairs.
- Romanian Instruct SFT — Romanian instruction fine-tuning data.
Research
Neurograf v1: An End-to-End Romanian-First Language Model Family
The Neurograf v1 technical report describes the model family, architecture, tokenizer, training methodology, and empirical scaling results.
Open Research
Neurograf is built around a simple objective: make Romanian-language AI research more accessible, reproducible, and open.
The project supports independent experimentation with Romanian language models, including pretraining, fine-tuning, evaluation, and language-resource development.
Model weights and supporting artifacts are published through the Neurograf Hugging Face organization.