Neurograf-6B-Instruct

A 6-billion-parameter Romanian-first instruction-tuned language model, trained from scratch and developed as part of the Neurograf research initiative.

Neurograf is an open research initiative dedicated to developing Romanian language models from the ground up, including tokenization, pretraining, instruction tuning, and inference.

Neurograf-6B-Instruct builds upon Neurograf-6B-Base, adapting the pretrained model for instruction following and conversational interaction.

Unlike multilingual models adapted to Romanian, the Neurograf family was pretrained primarily on Romanian text using a dedicated Romanian-first tokenizer.

Model architecture

Property Value
Architecture Modified NanoChat Transformer
Parameters Approximately 6B
Transformer layers 30
Hidden dimension 3,840
Attention heads 30
KV heads 30
Head dimension 128
Context length 2,048 tokens
Vocabulary 80,192 tokens
Tokenizer Custom Romanian-first BPE
Activation ReLU²
Normalization RMSNorm, QK normalization
Positional embeddings Rotary (RoPE)
Base model Neurograf-6B-Base

The underlying base model was pretrained on approximately 131 billion tokens, primarily from Romanian-language sources.

The model is distributed in Hugging Face Transformers format using sharded Safetensors weights.

Inference

Two inference implementations are included in this repository:

Script Backend Recommended use
infer_transformers.py Hugging Face Transformers Simple local inference and experimentation
infer_vllm.py vLLM GPU serving and high-throughput inference

Both examples use Neurograf's native conversation token format.

Option 1 — Hugging Face Transformers

Install the dependencies:

pip install "transformers>=5.3.0" torch accelerate

Download the repository:

git lfs install
git clone https://huggingface.co/Neurograf/Neurograf-6B-Instruct
cd Neurograf-6B-Instruct

Run the supplied inference script:

python infer_transformers.py

The model uses the NanoChatForCausalLM architecture, supported natively by Transformers 5.3.0.

No custom model implementation or trust_remote_code=True is required.

Option 2 — vLLM inference server

For GPU-accelerated serving, Neurograf can be used with vLLM.

Install the dependencies in a compatible CUDA environment:

pip install -U vllm requests

Step 1 — Start the vLLM server

CUDA_VISIBLE_DEVICES=0 vllm serve Neurograf/Neurograf-6B-Instruct \
  --host 127.0.0.1 \
  --port 8000 \
  --tensor-parallel-size 1 \
  --gpu-memory-utilization 0.95 \
  --max-model-len 2048 \
  --max-num-batched-tokens 1024 \
  > vllm_gpu0.log 2>&1 &

This configuration uses:

  • One GPU (CUDA_VISIBLE_DEVICES=0)
  • Tensor parallelism of 1
  • Up to 95% of available GPU memory
  • A maximum context length of 2,048 tokens
  • A maximum of 1,024 batched tokens per scheduling iteration
  • A local API server on port 8000
  • Background execution with logs saved to vllm_gpu0.log

GPU memory requirements depend on the precision, runtime overhead, and workload. A GPU with sufficient memory for the model weights and KV cache is required.

Step 2 — Run the inference client

In another terminal, from the downloaded repository:

python infer_vllm.py --model .

Or submit a single prompt:

python infer_vllm.py \
  --model . \
  --prompt "Explică-mi ce este fuziunea nucleară."

The inference client connects to the local vLLM server through its OpenAI-compatible completions API.

The server must already be running before launching the client.

Native conversation format

Neurograf uses dedicated tokens for conversation structure:

Token Token ID
<|bos|> 80174
<|user_start|> 80175
<|user_end|> 80176
<|assistant_start|> 80177
<|assistant_end|> 80178
<|system_start|> 80179
<|system_end|> 80180
<|turn_end|> 80191

The provided inference scripts construct prompts using the model's native token IDs, without requiring a Hugging Face chat template.

The examples are designed for straightforward, single-turn inference.

Research context

Neurograf-6B-Instruct belongs to the Neurograf v1 family of Romanian-first language models.

The base-model family spans five parameter scales:

All five base models share the same tokenizer and architectural design, enabling controlled studies of model scaling in Romanian.

Limitations

Neurograf-6B-Instruct is an experimental research model.

  • It can produce inaccurate, incomplete, or fabricated information.
  • Instruction following may vary depending on prompt structure and task complexity.
  • The context window is limited to 2,048 tokens.
  • The model is primarily intended for Romanian-language use.
  • Outputs should be independently verified in scientific, medical, legal, financial, or other consequential applications.

The release is intended to support open experimentation and further research on Romanian language modeling.

Project

Neurograf — Developing monolingual Romanian language models.

Explore the complete project and related releases:

Hugging Face — Neurograf


Neurograf is an independent, open research initiative focused on advancing Romanian-language artificial intelligence.

Downloads last month
26
Safetensors
Model size
6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Neurograf/Neurograf-6B-Instruct

Finetuned
(2)
this model

Collection including Neurograf/Neurograf-6B-Instruct