Bobic 1.5 Raye (125.8M)

Bobic 1.5 Raye is the flagship Small Language Model (SLM) in the Bobic series, featuring 125.86 million parameters. Designed with modern transformer optimizations (Grouped-Query Attention, SwiGLU, and NumberHead arithmetic injection), Bobic 1.5 Raye delivers strong reasoning and factual consistency while suppressing dialogue hallucinations.


Architectural Highlights

  • Grouped-Query Attention (GQA): 12 Query heads paired with 4 Key-Value heads (3:1 ratio). Dramatically reduces KV-cache latency and memory overhead while maintaining multi-head expressivity.
  • NumberHead Digit Grounding: Custom intra-number positional adapter that injects magnitude and place-value vectors directly into digit token hidden states, eliminating multi-digit arithmetic confusion.
  • SwiGLU Non-Linearity: 2048-dim SwiGLU feed-forward network provides dense representation capacity for scientific and dialogue tokens.
  • Targeted Alignment: Post-trained with targeted anti-hallucination and factual alignment datasets, cutting dialogue hallucination rates down to 21.4%.

Technical Specifications

Parameter Specification
Model Size 125,864,448 (125.86M parameters)
Layers (n_layers) 16
Hidden Size (d_model) 768
Attention Scheme GQA (12 Q-heads, 4 KV-heads, hd=64)
FFN Intermediate Size 2048 (SwiGLU)
Positional Embeddings Rotary Position Embeddings (RoPE, base=10,000)
Normalization RMSNorm
Head Projection Untied Head (vocab=16,384)
Precision bfloat16 / float32

Benchmark Performance

1. Academic Benchmark (MMLU-Pro)

Evaluated across 14 scientific disciplines on a 600-question sample with 10 options (A–J, random baseline ~10.00%):

  • Overall MMLU-Pro Accuracy: 20.00% (120 / 600)
    • Biology: 43.3%
    • Psychology: 40.0%
    • Economics: 28.6%
    • Engineering: 27.3%
    • History: 22.7%
    • Law: 21.9%
    • Philosophy: 18.2%
    • Chemistry: 17.6%
    • Physics: 12.5%

2. Dialogue & Hallucination Benchmark (bench_hallucinations.py)

  • Overall Consistency Score: 78.6%
  • Hallucination / Failure Rate: 21.4% (reduced from 71.4%)
  • Grounding Accuracy:
    • Arithmetic Grounding (2+2=4, 5+5=10, 10-4=6): 100.0%
    • Common Sense & Facts (ice is cold, summer grass is green): 75.0%
    • Identity & Role Preservation: 100.0%
    • False Claim Rejection (2+2=5 is false, elephants cannot fly): Verified

Quickstart (PyTorch)

import torch
from tokenizers import Tokenizer, decoders
from model import Bobic15Raye, CFG_RAYE

dev = torch.device('cuda' if torch.cuda.is_available() else 'cpu')

# 1. Load Tokenizer
tok = Tokenizer.from_file('tokenizer.json')
tok.decoder = decoders.ByteLevel()
digit_ids = [tok.encode(str(d)).ids[0] for d in range(10)]

# 2. Instantiate Model and Load Aligned Weights
model = Bobic15Raye(cfg=CFG_RAYE, digit_ids=digit_ids).to(dev)
ck = torch.load('bobic15_raye.pt', map_location=dev)
model.load_state_dict(ck['model'])
model.eval()

# 3. Text Generation
prompt = "User: привет, как дела?\nBobic:"
ids = tok.encode(prompt, add_special_tokens=False).ids
idx = torch.tensor([ids], device=dev)

with torch.no_grad():
    out = model.generate(idx, n=35, temp=0.4)

output_text = tok.decode(out[0].tolist())
print(output_text[len(prompt):].strip().split('\n')[0])

License

Released under the MIT License. Created by Skebobic.

Downloads last month
16
GGUF
Model size
0.1B params
Architecture
bobic
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results