Usaid AI (500M) โ Causal Language Model
Created and Developed by: Mohamed Usaid
Architecture: Decoder-only Transformer with GQA, RoPE, SwiGLU, and RMSNorm
Model Size: 500,136,960 parameters
Weight Format: FP16 SafeTensors
Repository: UsaidAI-500M
Model Summary
Usaid AI (500M) is an experimental open-weights causal language model developed and trained from scratch by Mohamed Usaid.
The model uses a modern decoder-only Transformer architecture incorporating Grouped-Query Attention (GQA), Rotary Position Embeddings (RoPE), SwiGLU activations, and RMSNorm.
The model contains 500,136,960 parameters and was pretrained for approximately 131 million tokens using cloud GPU infrastructure (dual NVIDIA Tesla T4 GPUs). It was subsequently fine-tuned using supervised instruction and conversational data.
This release represents an experimental milestone in building and training a language model from scratch under constrained compute resources.
Training Status
The model is substantially undertrained relative to its parameter count. It was trained on approximately 131M pretraining tokens, corresponding to approximately 0.26 training tokens per parameter.
For comparison, Chinchilla-style compute-optimal scaling studies found approximately 20 training tokens per parameter at the compute budgets studied. This should be treated as a scaling-law reference rather than a hard threshold for model capability.
Consequently, this model should be viewed primarily as an experimental research and engineering artifact rather than a production-grade general-purpose language model.
Architecture Specifications
The model incorporates modern decoder-only Transformer components:
- Model Scale: Half-billion parameter decoder-only model (500,136,960 parameters)
- Transformer Layers: 30
- Hidden Dimension (d_model): 1024
- Query Attention Heads (n_q): 16
- Key/Value Attention Heads (n_kv): 4 (4:1 Grouped-Query Attention)
- Attention Head Dimension (d_k): 64
- Feed-Forward Network: SwiGLU (intermediate size: 3456)
- Positional Embeddings: Rotary Position Embeddings (RoPE)
- Normalization: RMSNorm (eps = 1e-5)
- Weight Tying: Untied input embedding and output projection heads
- Tokenizer: Byte-Pair Encoding (GPT-2 BPE, 50,257 vocabulary)
- Maximum Context Capacity: 2,048 tokens
Training Details
1. Pretraining
- Compute Environment: Cloud GPU infrastructure (2ร NVIDIA Tesla T4)
- Pretraining Tokens: 131,072,000 tokens (FineWeb-Edu subset)
- Training Steps: 2,000 steps
- Effective Global Batch: 64 sequences ร 1,024 tokens (65,536 tokens/step)
- Training Duration: Approximately 7.33 hours wall-clock
- Validation Loss: 2.9091
- Validation Perplexity: 18.34
2. Supervised Fine-Tuning (SFT)
- Compute Environment: Cloud GPU infrastructure (2ร NVIDIA Tesla T4)
- Dataset: 4,731 multi-turn conversational and identity instruction pairs
- Training Steps: 441 steps
- Epochs: 3
- Initial Loss: 2.9302
- Best SFT Loss: 1.3938
The SFT stage substantially reduced training loss and aligned conversational behavior, but the resulting model should still be considered experimental given the limited pretraining token budget.
Evaluation
A targeted 20-question diagnostic probe was used to compare the pretrained base model and the SFT-aligned model using length-normalized completion log-likelihood scoring (random chance baseline: 25%):
| Benchmark Category | Pre-SFT Base Model | Post-SFT Aligned Model | Empirical Delta |
|---|---|---|---|
| ML & Transformer Architecture | 2 / 5 (40%) | 3 / 5 (60%) | +20% (+1 item) |
| World Knowledge & Science | 2 / 5 (40%) | 3 / 5 (60%) | +20% (+1 item) |
| Python & Software Engineering | 1 / 5 (20%) | 1 / 5 (20%) | 0% (baseline) |
| Logic & Arithmetic | 1 / 5 (20%) | 1 / 5 (20%) | 0% (baseline) |
| OVERALL DIAGNOSTIC ACCURACY | 6 / 20 (30%) | 8 / 20 (40%) | +10% (+2 items) |
This 20-question probe is a small diagnostic evaluation and should not be interpreted as a standardized benchmark. The model shows measurable performance on the included diagnostic probe, but broader capabilities remain bounded by the pretraining scale.
Known Limitations
The model was trained on approximately 131M tokens (0.26 training tokens per parameter). This is substantially below the approximately 20 tokens-per-parameter ratio commonly associated with Chinchilla-style compute-optimal training at the studied compute budgets. This ratio should be treated as a scaling-law reference rather than a hard threshold for reasoning ability.
As a result, the model is substantially undertrained relative to its parameter count. Complex multi-step code generation, arithmetic, reasoning, and factual recall may be unreliable.
Other observed limitations include:
- Unreliable multi-step reasoning
- Arithmetic and symbolic calculation errors
- Hallucinated factual statements
- Repetition loops in extended generations
- Sensitivity to prompt formatting and generation parameters
The model is therefore intended primarily for experimentation, research, architecture studies, and educational purposes.
Quickstart & Usage
1. Native TinyGPT Repository (Primary Reference Implementation)
# Clone the repository
git clone -b scaling-up-500m https://github.com/mohamedusaid/TinyGPT.git
cd TinyGPT
# Continuous interactive multi-turn chat
python scripts/run_chat_loop.py
# Continuous chat with RAG knowledge retrieval enabled
python scripts/run_chat_loop.py --rag
2. Loading with Hugging Face Transformers
The exported checkpoint is packaged for compatibility with Hugging Face's LlamaForCausalLM implementation:
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "Usaidddddddddddddd/UsaidAI-500M"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto"
)
prompt = "User: What is the capital of France?\nAssistant:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
output_tokens = model.generate(
**inputs,
max_new_tokens=60,
temperature=0.7,
top_p=0.9,
repetition_penalty=1.15,
eos_token_id=tokenizer.eos_token_id
)
print(tokenizer.decode(output_tokens[0], skip_special_tokens=True))
Citation & Attribution
If you use Usaid AI in your research or project, please credit:
@misc{usaid_ai_2026,
author = {Mohamed Usaid},
title = {Usaid AI: A 500M Parameter Modern Causal Language Model with Grouped-Query Attention},
year = {2026},
publisher = {Hugging Face / GitHub},
howpublished = {\url{https://huggingface.co/Usaidddddddddddddd/UsaidAI-500M}}
}
- Downloads last month
- 3