Neurograf-560M-Base

Neurograf-560M-Base is an approximately 560-million-parameter decoder-only language model pretrained from scratch on 131 billion Romanian-language tokens.

It belongs to Neurograf v1, an open research initiative focused on developing Romanian-first language models, tokenizers, and training resources.

This is a base language model, not an instruction-tuned assistant. It is intended for text continuation, language modeling research, and further fine-tuning.

Model Architecture

Neurograf uses a modified Nanochat transformer architecture with rotary positional embeddings, QK normalization, parameter-free RMSNorm, ReLU² feed-forward activations, and untied input/output embeddings.

Parameter Value
Architecture Decoder-only Transformer
Parameters Approximately 560M
Transformer layers 12
Hidden dimension 1,536
Attention heads 12
Key/value heads 12
Head dimension 128
Vocabulary size 80,192
Training sequence length 2,048
Positional encoding RoPE
Attention Multi-head causal attention
Activation ReLU²
Normalization RMSNorm, QK normalization
Output logits Tanh softcap (15)
Training framework Modified Nanochat
Checkpoint Step 250,000

The architecture supports grouped-query attention, although this model uses standard multi-head attention with equal numbers of query and key/value heads.

Training Data

Neurograf-560M-Base was pretrained on approximately 131 billion tokens of Romanian-language text, using the Neurograf tokenizer developed specifically for the project.

The training corpus contains approximately 170 million Romanian text blocks collected and processed for language modeling.

All five Neurograf v1 base models were trained using the same overall token budget, enabling direct comparisons across model sizes.

Pretraining Results

Training Diagnostics

The following figure presents training loss, gradient norms, and validation bits per byte (BPB) for the five Neurograf v1 base models.

Neurograf pretraining diagnostics

Across the model family, larger models generally achieve lower validation BPB, with continued improvement throughout pretraining.

The final 560M checkpoint achieves:

Validation BPB: 0.742747

The checkpoint was saved at training step 250,000, corresponding to approximately 131.1 billion training-token positions.

This checkpoint also achieves the lowest recorded validation BPB for its training run, as reported in the checkpoint metadata.

Fixed-Data Scaling Law

We examine how validation performance scales with model size when the training token budget is held fixed at approximately 131 billion tokens.

Neurograf fixed-data scaling law

The empirical relationship is described by:

BPB(N) = 0.437 + 144 × N^(-0.305)

with a reported fit of R² = 0.99999 across the five model sizes.

This relationship describes the observed Neurograf v1 results under a fixed training-data budget. Extrapolations beyond the measured model sizes should be treated as exploratory rather than established performance predictions.

Model Files

The repository contains the original PyTorch checkpoint and tokenizer artifacts:

File Description
model_250000.pt Final pretrained model weights
meta_250000.json Architecture configuration and training metadata
tokenizer.pkl Original tiktoken-based BPE tokenizer
token_bytes.pt Token-byte mapping artifact

The model is distributed in its original Nanochat-compatible checkpoint format. It has not yet been converted into a standard Hugging Face Transformers model.

Loading requires a compatible implementation of the modified Neurograf/Nanochat architecture. Inference examples and additional compatibility tools may be published separately.

Security note: The original .pt and .pkl formats use Python serialization mechanisms. Load these files only from trusted sources.

Intended Uses

  • Romanian-language modeling research
  • Continued pretraining
  • Supervised fine-tuning
  • Evaluation of Romanian language capabilities
  • Model scaling and training dynamics research
  • Development of specialized Romanian-language models

Limitations

This is a pretrained base model and has not been optimized for conversational instruction following.

Generated text may contain factual inaccuracies, biases, repetitions, or inappropriate content. The model has not been comprehensively evaluated for safety or reliability in high-stakes applications.

The training context length is 2,048 tokens. Longer-context behavior has not been systematically validated.

License

Released under the Apache License 2.0, subject to applicable third-party rights and attribution requirements.

Project

Neurograf — Developing monolingual Romanian language models from scratch.

Neurograf on Hugging Face

Citation

If you use Neurograf in your research, please cite the Neurograf v1 technical report:

Neurograf v1: An End-to-End Romanian-First Language Model Family.

Full bibliographic information and a BibTeX citation can be added here.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including Neurograf/Neurograf-560M-Base