Instructions to use Neurograf/Neurograf-6B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Neurograf/Neurograf-6B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Neurograf/Neurograf-6B-Instruct")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Neurograf/Neurograf-6B-Instruct") model = AutoModelForCausalLM.from_pretrained("Neurograf/Neurograf-6B-Instruct", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Neurograf/Neurograf-6B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Neurograf/Neurograf-6B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Neurograf/Neurograf-6B-Instruct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Neurograf/Neurograf-6B-Instruct
- SGLang
How to use Neurograf/Neurograf-6B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Neurograf/Neurograf-6B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Neurograf/Neurograf-6B-Instruct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Neurograf/Neurograf-6B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Neurograf/Neurograf-6B-Instruct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Neurograf/Neurograf-6B-Instruct with Docker Model Runner:
docker model run hf.co/Neurograf/Neurograf-6B-Instruct
Neurograf-6B-Instruct
A 6-billion-parameter Romanian-first instruction-tuned language model, trained from scratch and developed as part of the Neurograf research initiative.
Neurograf is an open research initiative dedicated to developing Romanian language models from the ground up, including tokenization, pretraining, instruction tuning, and inference.
Neurograf-6B-Instruct builds upon Neurograf-6B-Base, adapting the pretrained model for instruction following and conversational interaction.
Unlike multilingual models adapted to Romanian, the Neurograf family was pretrained primarily on Romanian text using a dedicated Romanian-first tokenizer.
Model architecture
| Property | Value |
|---|---|
| Architecture | Modified NanoChat Transformer |
| Parameters | Approximately 6B |
| Transformer layers | 30 |
| Hidden dimension | 3,840 |
| Attention heads | 30 |
| KV heads | 30 |
| Head dimension | 128 |
| Context length | 2,048 tokens |
| Vocabulary | 80,192 tokens |
| Tokenizer | Custom Romanian-first BPE |
| Activation | ReLU² |
| Normalization | RMSNorm, QK normalization |
| Positional embeddings | Rotary (RoPE) |
| Base model | Neurograf-6B-Base |
The underlying base model was pretrained on approximately 131 billion tokens, primarily from Romanian-language sources.
The model is distributed in Hugging Face Transformers format using sharded Safetensors weights.
Inference
Two inference implementations are included in this repository:
| Script | Backend | Recommended use |
|---|---|---|
infer_transformers.py |
Hugging Face Transformers | Simple local inference and experimentation |
infer_vllm.py |
vLLM | GPU serving and high-throughput inference |
Both examples use Neurograf's native conversation token format.
Option 1 — Hugging Face Transformers
Install the dependencies:
pip install "transformers>=5.3.0" torch accelerate
Download the repository:
git lfs install
git clone https://huggingface.co/Neurograf/Neurograf-6B-Instruct
cd Neurograf-6B-Instruct
Run the supplied inference script:
python infer_transformers.py
The model uses the NanoChatForCausalLM architecture, supported natively by Transformers 5.3.0.
No custom model implementation or trust_remote_code=True is required.
Option 2 — vLLM inference server
For GPU-accelerated serving, Neurograf can be used with vLLM.
Install the dependencies in a compatible CUDA environment:
pip install -U vllm requests
Step 1 — Start the vLLM server
CUDA_VISIBLE_DEVICES=0 vllm serve Neurograf/Neurograf-6B-Instruct \
--host 127.0.0.1 \
--port 8000 \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.95 \
--max-model-len 2048 \
--max-num-batched-tokens 1024 \
> vllm_gpu0.log 2>&1 &
This configuration uses:
- One GPU (
CUDA_VISIBLE_DEVICES=0) - Tensor parallelism of 1
- Up to 95% of available GPU memory
- A maximum context length of 2,048 tokens
- A maximum of 1,024 batched tokens per scheduling iteration
- A local API server on port 8000
- Background execution with logs saved to
vllm_gpu0.log
GPU memory requirements depend on the precision, runtime overhead, and workload. A GPU with sufficient memory for the model weights and KV cache is required.
Step 2 — Run the inference client
In another terminal, from the downloaded repository:
python infer_vllm.py --model .
Or submit a single prompt:
python infer_vllm.py \
--model . \
--prompt "Explică-mi ce este fuziunea nucleară."
The inference client connects to the local vLLM server through its OpenAI-compatible completions API.
The server must already be running before launching the client.
Native conversation format
Neurograf uses dedicated tokens for conversation structure:
| Token | Token ID |
|---|---|
<|bos|> |
80174 |
<|user_start|> |
80175 |
<|user_end|> |
80176 |
<|assistant_start|> |
80177 |
<|assistant_end|> |
80178 |
<|system_start|> |
80179 |
<|system_end|> |
80180 |
<|turn_end|> |
80191 |
The provided inference scripts construct prompts using the model's native token IDs, without requiring a Hugging Face chat template.
The examples are designed for straightforward, single-turn inference.
Research context
Neurograf-6B-Instruct belongs to the Neurograf v1 family of Romanian-first language models.
The base-model family spans five parameter scales:
| Model | Parameters |
|---|---|
| Neurograf-160M-Base | 160M |
| Neurograf-560M-Base | 560M |
| Neurograf-1.5B-Base | 1.5B |
| Neurograf-3.2B-Base | 3.2B |
| Neurograf-6B-Base | 6B |
All five base models share the same tokenizer and architectural design, enabling controlled studies of model scaling in Romanian.
Limitations
Neurograf-6B-Instruct is an experimental research model.
- It can produce inaccurate, incomplete, or fabricated information.
- Instruction following may vary depending on prompt structure and task complexity.
- The context window is limited to 2,048 tokens.
- The model is primarily intended for Romanian-language use.
- Outputs should be independently verified in scientific, medical, legal, financial, or other consequential applications.
The release is intended to support open experimentation and further research on Romanian language modeling.
Project
Neurograf — Developing monolingual Romanian language models.
Explore the complete project and related releases:
Neurograf is an independent, open research initiative focused on advancing Romanian-language artificial intelligence.
- Downloads last month
- 26
Model tree for Neurograf/Neurograf-6B-Instruct
Base model
Neurograf/Neurograf-6B-Base