Instructions to use Skebobic/Bobic-1.5-Lean with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Skebobic/Bobic-1.5-Lean with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Skebobic/Bobic-1.5-Lean:Q8_0 # Run inference directly in the terminal: llama cli -hf Skebobic/Bobic-1.5-Lean:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Skebobic/Bobic-1.5-Lean:Q8_0 # Run inference directly in the terminal: llama cli -hf Skebobic/Bobic-1.5-Lean:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Skebobic/Bobic-1.5-Lean:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf Skebobic/Bobic-1.5-Lean:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Skebobic/Bobic-1.5-Lean:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Skebobic/Bobic-1.5-Lean:Q8_0
Use Docker
docker model run hf.co/Skebobic/Bobic-1.5-Lean:Q8_0
- LM Studio
- Jan
- vLLM
How to use Skebobic/Bobic-1.5-Lean with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Skebobic/Bobic-1.5-Lean" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Skebobic/Bobic-1.5-Lean", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Skebobic/Bobic-1.5-Lean:Q8_0
- Ollama
How to use Skebobic/Bobic-1.5-Lean with Ollama:
ollama run hf.co/Skebobic/Bobic-1.5-Lean:Q8_0
- Unsloth Desktop
- Docker Model Runner
How to use Skebobic/Bobic-1.5-Lean with Docker Model Runner:
docker model run hf.co/Skebobic/Bobic-1.5-Lean:Q8_0
- Lemonade
How to use Skebobic/Bobic-1.5-Lean with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Skebobic/Bobic-1.5-Lean:Q8_0
Run and chat with the model
lemonade run user.Bobic-1.5-Lean-Q8_0
List all available models
lemonade list
- Atomic Chat
Bobic 1.5 Lean (60.8M)
Bobic 1.5 Lean is an ultra-compact Small Language Model (SLM) with 60.8 million parameters, optimized for micro-scale edge deployment, mathematical arithmetic, and fast conversational inference.
Trained on a curated multi-source corpus including Wikipedia, Telegram conversation logs, high-density scientific QA (physics, arithmetic, trivia), and academic reasoning splits (ARC, OpenBookQA).
Model Architecture & Specifications
| Parameter | Specification | Notes |
|---|---|---|
| Total Parameters | 60,840,960 (60.8M) | Deep & narrow design |
Layers (n_layers) |
16 | Maximizes depth over width at small scale |
Hidden Dimension (d_model) |
512 | Ultra-low memory footprint |
| Attention Mechanism | MQA (Multi-Query Attention) | 8 query heads (hd=64), 1 shared KV head |
| Feed-Forward Network (FFN) | SwiGLU | Hidden dimension = 1408 (~2.75x expansion) |
| Positional Embeddings | RoPE (Rotary Position Embeddings) | Base frequency = 10,000 |
| Normalization | RMSNorm | Pre-layer normalization with learned scaling |
| Vocabulary | 16,384 tokens | ByteLevel BPE tokenizer (bobic15_tok.json) |
| NumberHead Adapter | Integrated | Intra-number position embedding injection for digits |
Key Benchmark Results
- MMLU-Pro Benchmark (600q subset): 19.83% (Random baseline with 10 options A–J is ~10.00%).
- Memory Footprint: Less than 1.0 GB VRAM during full sequence inference.
- Throughput: Over 40,000 tokens/sec training speed on a single NVIDIA GeForce RTX 4070.
Quickstart & Usage (PyTorch)
import torch
from tokenizers import Tokenizer, decoders
from model import Bobic15Lean, CFG
dev = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
# 1. Load Tokenizer
tok = Tokenizer.from_file('tokenizer.json')
tok.decoder = decoders.ByteLevel()
digit_ids = [tok.encode(str(d)).ids[0] for d in range(10)]
# 2. Initialize Model and Load Weights
model = Bobic15Lean(cfg=CFG, digit_ids=digit_ids).to(dev)
ck = torch.load('bobic15_lean.pt', map_location=dev)
model.load_state_dict(ck['model'])
model.eval()
# 3. Generate Response
prompt = "User: привет\nBobic:"
ids = tok.encode(prompt, add_special_tokens=False).ids
idx = torch.tensor([ids], device=dev)
with torch.no_grad():
out = model.generate(idx, n=30, temp=0.4)
output_text = tok.decode(out[0].tolist())
print(output_text[len(prompt):].strip().split('\n')[0])
License & Attribution
Released under the MIT License. Created and trained by Skebobic.
- Downloads last month
- 17
Hardware compatibility
Log In to add your hardware
8-bit
Evaluation results
- Accuracy on MMLU-Proself-reported19.830