Instructions to use Skebobic/Bobic-1.5-Raye with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Skebobic/Bobic-1.5-Raye with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Skebobic/Bobic-1.5-Raye:Q8_0 # Run inference directly in the terminal: llama cli -hf Skebobic/Bobic-1.5-Raye:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Skebobic/Bobic-1.5-Raye:Q8_0 # Run inference directly in the terminal: llama cli -hf Skebobic/Bobic-1.5-Raye:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Skebobic/Bobic-1.5-Raye:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf Skebobic/Bobic-1.5-Raye:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Skebobic/Bobic-1.5-Raye:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Skebobic/Bobic-1.5-Raye:Q8_0
Use Docker
docker model run hf.co/Skebobic/Bobic-1.5-Raye:Q8_0
- LM Studio
- Jan
- vLLM
How to use Skebobic/Bobic-1.5-Raye with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Skebobic/Bobic-1.5-Raye" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Skebobic/Bobic-1.5-Raye", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Skebobic/Bobic-1.5-Raye:Q8_0
- Ollama
How to use Skebobic/Bobic-1.5-Raye with Ollama:
ollama run hf.co/Skebobic/Bobic-1.5-Raye:Q8_0
- Unsloth Desktop
- Docker Model Runner
How to use Skebobic/Bobic-1.5-Raye with Docker Model Runner:
docker model run hf.co/Skebobic/Bobic-1.5-Raye:Q8_0
- Lemonade
How to use Skebobic/Bobic-1.5-Raye with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Skebobic/Bobic-1.5-Raye:Q8_0
Run and chat with the model
lemonade run user.Bobic-1.5-Raye-Q8_0
List all available models
lemonade list
- Atomic Chat
Bobic 1.5 Raye (125.8M)
Bobic 1.5 Raye is the flagship Small Language Model (SLM) in the Bobic series, featuring 125.86 million parameters. Designed with modern transformer optimizations (Grouped-Query Attention, SwiGLU, and NumberHead arithmetic injection), Bobic 1.5 Raye delivers strong reasoning and factual consistency while suppressing dialogue hallucinations.
Architectural Highlights
- Grouped-Query Attention (GQA): 12 Query heads paired with 4 Key-Value heads (3:1 ratio). Dramatically reduces KV-cache latency and memory overhead while maintaining multi-head expressivity.
- NumberHead Digit Grounding: Custom intra-number positional adapter that injects magnitude and place-value vectors directly into digit token hidden states, eliminating multi-digit arithmetic confusion.
- SwiGLU Non-Linearity: 2048-dim SwiGLU feed-forward network provides dense representation capacity for scientific and dialogue tokens.
- Targeted Alignment: Post-trained with targeted anti-hallucination and factual alignment datasets, cutting dialogue hallucination rates down to 21.4%.
Technical Specifications
| Parameter | Specification |
|---|---|
| Model Size | 125,864,448 (125.86M parameters) |
Layers (n_layers) |
16 |
Hidden Size (d_model) |
768 |
| Attention Scheme | GQA (12 Q-heads, 4 KV-heads, hd=64) |
| FFN Intermediate Size | 2048 (SwiGLU) |
| Positional Embeddings | Rotary Position Embeddings (RoPE, base=10,000) |
| Normalization | RMSNorm |
| Head Projection | Untied Head (vocab=16,384) |
| Precision | bfloat16 / float32 |
Benchmark Performance
1. Academic Benchmark (MMLU-Pro)
Evaluated across 14 scientific disciplines on a 600-question sample with 10 options (A–J, random baseline ~10.00%):
- Overall MMLU-Pro Accuracy: 20.00% (120 / 600)
- Biology: 43.3%
- Psychology: 40.0%
- Economics: 28.6%
- Engineering: 27.3%
- History: 22.7%
- Law: 21.9%
- Philosophy: 18.2%
- Chemistry: 17.6%
- Physics: 12.5%
2. Dialogue & Hallucination Benchmark (bench_hallucinations.py)
- Overall Consistency Score: 78.6%
- Hallucination / Failure Rate: 21.4% (reduced from 71.4%)
- Grounding Accuracy:
- Arithmetic Grounding (
2+2=4,5+5=10,10-4=6): 100.0% - Common Sense & Facts (ice is cold, summer grass is green): 75.0%
- Identity & Role Preservation: 100.0%
- False Claim Rejection (
2+2=5 is false,elephants cannot fly): Verified
- Arithmetic Grounding (
Quickstart (PyTorch)
import torch
from tokenizers import Tokenizer, decoders
from model import Bobic15Raye, CFG_RAYE
dev = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
# 1. Load Tokenizer
tok = Tokenizer.from_file('tokenizer.json')
tok.decoder = decoders.ByteLevel()
digit_ids = [tok.encode(str(d)).ids[0] for d in range(10)]
# 2. Instantiate Model and Load Aligned Weights
model = Bobic15Raye(cfg=CFG_RAYE, digit_ids=digit_ids).to(dev)
ck = torch.load('bobic15_raye.pt', map_location=dev)
model.load_state_dict(ck['model'])
model.eval()
# 3. Text Generation
prompt = "User: привет, как дела?\nBobic:"
ids = tok.encode(prompt, add_special_tokens=False).ids
idx = torch.tensor([ids], device=dev)
with torch.no_grad():
out = model.generate(idx, n=35, temp=0.4)
output_text = tok.decode(out[0].tolist())
print(output_text[len(prompt):].strip().split('\n')[0])
License
Released under the MIT License. Created by Skebobic.
- Downloads last month
- 16
8-bit
Evaluation results
- Accuracy on MMLU-Proself-reported20.000