G0-nano-instruct

A 62M-parameter GPT pretrained and instruction-tuned entirely on a single 8GB-RAM NVIDIA Jetson device, without cloud infrastructure or multi-GPU setups. Chat-oriented checkpoint for short instruction-following interactions.

Overview

G0-nano-instruct is the instruction-tuned version of the G0 Nano model. Both pretraining and supervised fine-tuning were performed under an 8GB unified-memory constraint.

The model is intentionally small. Its purpose is to demonstrate a complete from-scratch training workflow on modest hardware, not to compete with much larger language models on broad factual knowledge.

Model variants

The raw pretrained version of the same model is available as G0-nano-base.

What this version adds

Compared with G0-nano-base, this checkpoint adds supervised instruction fine-tuning and a chat format. It is designed for short, single-turn interactions rather than long conversations.

Architecture

Llama-style decoder-only Transformer:

Property Value
Parameters 62.1M, with embeddings shared with the language-model head
Layers 12
Hidden size 640
Attention Grouped-Query Attention, 10 query heads / 2 key-value heads, head dimension 64
Position encoding RoPE, θ=10000
Feed-forward network SwiGLU, hidden dimension 1728
Normalization RMSNorm
Context length 1024 tokens; 512 tokens during supervised fine-tuning
Vocabulary 16,388 tokens: 16,384 SentencePiece tokens plus 4 chat tokens

Training

  • Pretraining data: approximately 1.5B tokens of English web and book text
  • Sources: FineWeb-Edu, BookCorpus, OpenWebText, PG-19 and WikiHow
  • Instruction tuning: cleaned Alpaca instruction/response data
  • Objective: causal next-token prediction followed by supervised instruction fine-tuning
  • Training hardware: a single NVIDIA Jetson with 8GB of unified memory

Usage

Hugging Face Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "AZERDSQ/G0-nano-instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)

inputs = tokenizer("What is the capital of France?", return_tensors="pt")
outputs = model.generate(
    **inputs,
    max_new_tokens=50,
    do_sample=True,
    top_k=50,
    temperature=0.8,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

trust_remote_code=True is required because this repository uses a custom Transformer implementation.

Ollama

ollama run azerdsq/g0-nano-instruct

The chat format is built into the model for simple instruction-following use.

Benchmarks

Limitations

  • 62M parameters impose a hard limit on factual knowledge; expect fluent but frequently incorrect answers on knowledge-intensive prompts.
  • Maximum context length is 1024 tokens.
  • English-only training data.
  • Designed for short, single-turn exchanges; it is not a long-context conversational model.
  • Single-sequence generation only; padded batched inference is not supported by the custom model code.

This model should not be used for high-stakes decisions, factual verification, medical advice, legal advice or autonomous actions.

License

Apache 2.0. This release contains model weights and the code required to load them; it does not include the training data or private training infrastructure.

Links

Downloads last month
169
Safetensors
Model size
62.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support