Neurograf

community
Activity Feed

AI & ML interests

Developing monolingual Romanian models.

Recent Activity

rotarue  updated a collection 2 days ago
Neurograf v1 — SFT models
rotarue  updated a collection 2 days ago
Neurograf v1 — SFT models
rotarue  updated a model 2 days ago
Neurograf/Neurograf-6B-Translate
View all activity

Organization Card

Neurograf v1

Developing Romanian language models from scratch.

Neurograf is an open research initiative focused on training Romanian-first language models, developing language resources, and making AI research accessible to everyone.

Neurograf v1 Model Family

Neurograf v1 consists of five decoder-only Transformer models, ranging from 160 million to approximately 6 billion parameters.

All five models were trained from scratch on approximately 131 billion Romanian-language tokens, using a shared Romanian-first tokenizer and a modified Nanochat architecture.

Model Parameters Layers Validation BPB
Neurograf-160M-Base 160M 6 0.886683
Neurograf-560M-Base 560M 12 0.742747
Neurograf-1.5B-Base 1.5B 18 0.665701
Neurograf-3.2B-Base 3.2B 24 0.619530
Neurograf-6B-Base 6B 30 0.587616

All models use a context length of 2,048 tokens and a shared vocabulary of 80,192 tokens.

These are pretrained base models intended for research, continued pretraining, and downstream fine-tuning.

Pretraining Results

The Neurograf v1 models provide a controlled comparison of language-model scaling across five architectures trained with approximately the same token budget.

Training Diagnostics

The training results show consistent improvements in validation bits per byte (BPB) as model size increases.

Fixed-Data Scaling

A power-law fit to the five measured model sizes achieves R² = 0.99999, illustrating the relationship between model capacity and validation performance under a fixed Romanian-language training budget.

Extrapolations beyond the evaluated model sizes remain exploratory.

Architecture and Tokenization

The Neurograf v1 models share a modified Nanochat Transformer architecture featuring:

  • Rotary positional embeddings (RoPE)
  • QK normalization
  • RMSNorm
  • ReLU² feed-forward activations
  • Untied input and output embeddings
  • A custom Romanian-first BPE tokenizer

The models are currently distributed as original PyTorch checkpoints with their corresponding architecture metadata and tokenizer files.

Datasets and Language Resources

In addition to pretrained models, Neurograf develops datasets and resources for Romanian-language training and adaptation.

Research

Neurograf v1: An End-to-End Romanian-First Language Model Family

The Neurograf v1 technical report describes the model family, architecture, tokenizer, training methodology, and empirical scaling results.

Open Research

Neurograf is built around a simple objective: make Romanian-language AI research more accessible, reproducible, and open.

The project supports independent experimentation with Romanian language models, including pretraining, fine-tuning, evaluation, and language-resource development.

Model weights and supporting artifacts are published through the Neurograf Hugging Face organization.