Instructions to use harims95/hobbylm-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use harims95/hobbylm-1B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="harims95/hobbylm-1B", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("harims95/hobbylm-1B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use harims95/hobbylm-1B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "harims95/hobbylm-1B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harims95/hobbylm-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/harims95/hobbylm-1B
- SGLang
How to use harims95/hobbylm-1B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "harims95/hobbylm-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harims95/hobbylm-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "harims95/hobbylm-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harims95/hobbylm-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use harims95/hobbylm-1B with Docker Model Runner:
docker model run hf.co/harims95/hobbylm-1B
HobbyLM-1B
A 1.037B-parameter sparse mixture-of-experts language model trained from scratch on 100B tokens, with approximately 305M parameters active per token. This repository holds the final annealed base model.
HobbyLM by Fuel Labs
Instruct | GGUF | Checkpoints | Online Demo | GitHub | Windows App
Highlights
- Sparse MoE, trained from scratch: 1.037B total parameters, approximately 305M active per token (64 routed experts + 1 shared expert, top-8 routing), pretrained on 100B tokens.
- Open release: the final annealed base (this repository), the instruction-tuned model, the 4K context-extension checkpoint and GGUF files, all under Apache-2.0.
- Runs on an ordinary CPU: GGUF files, patched llama.cpp and Ollama runtimes, and a double-click chat app for Windows.
- Measured, with its limits stated: a release evaluation of the base and instruction-tuned models, plus function calling (BFCL) and instruction following (IFEval), with setup and provenance.
Model List
| Model | Format | Where | Notes |
|---|---|---|---|
| HobbyLM-1B (base) | Transformers, fp32 | this repository (root files) | final annealed base (step 95,367), 1024 context |
| HobbyLM-1B Instruct | Transformers, fp32 | hobbylm-1B-instruct |
Instruction-tuned, 4096 context; runs in the online demo |
| Checkpoints | Transformers, fp32 | hobbylm-1b-checkpoints |
base-anneal-final, context-4k-step150, sft-step3450 |
| GGUF | F32 / Q8_0 / Q4_K_M | hobbylm-1B-gguf |
Instruct in three precisions plus base F32; needs a patched runtime |
| Windows chat app and runtimes | ZIP | GitHub release v1.0.0 | double-click chat app; patched llama.cpp and Ollama for Windows x64 CPU |
Use the tested Quickstart instructions below. Hugging Face’s automatic ‘Use this model’ snippets are not validated for HobbyLM. Our GGUF release requires the supplied patched runtimes. Stock Ollama v0.35.1 failed our loading test; LM Studio, Docker Model Runner, vLLM and SGLang have not been validated.
Model Information
| Parameters | 1,037,121,536 total; 304,953,344 active per token (embeddings counted once; tied LM head) |
| Weights | model.safetensors, float32. Raw training checkpoint 1b_flagship/model.pt, SHA-256 947e80d903f3b2bf9e55af9eee785dfa49740b2d4b851bc564fc1e65e085f843 |
| Export identity | Verified (2026-10-06): all 279 stored tensors are bit-for-bit identical to the raw checkpoint after conversion. The tied lm_head is not stored separately. |
| Precision | fp32 only (see "Precision") |
| Tokenizer | GPT-2 BPE (50,257 tokens; embedding matrix padded to 50,304 rows). `< |
| Context | 1024 tokens (pretraining sequence length) |
| Component | Value |
|---|---|
| Layers | 20 (layer 0 dense SwiGLU, intermediate 2816; layers 1–19 MoE) |
| Experts | 64 routed (intermediate 224) + 1 shared; top-8 routed per token |
| Router | sigmoid gating with aux-loss-free bias balancing (DeepSeek-V3 style) |
| Attention | GQA, 16 query heads / 8 KV heads, head dim 128, per-head QK RMSNorm |
| Position | plain RoPE, theta 10000, no scaling |
| Hidden size / norm | 1024 / RMSNorm (eps 1e-6) |
| Embeddings | tied input embedding / LM head |
Introduction
HobbyLM-1B is a compact sparse mixture-of-experts language model developed by Fuel Labs. It was trained from scratch on 100B tokens to study how much capability can be developed within a comparatively small active-parameter budget.
The files at the root of this repository are the final annealed base model: the end of pretraining after the anneal (step 95,367). It continues text: it is not instruction-tuned, not preference-tuned and not safety-aligned. For chat, extraction, rewriting and single function calls, use HobbyLM-1B Instruct, or try it in the online demo.
Base-Model Benchmark Context
The following figures place selected HobbyLM Base results alongside public results reported for compact peer models. Scores are percentages and higher is better. Peer evaluations come from public model cards and technical reports; implementations may differ across sources. The source notes and scale treatment are included directly in each figure.
What Changed After Instruction Tuning
The instruction-tuned model (HobbyLM-1B Instruct) improves several release-evaluation results under the matched float32, zero-shot protocol shown below. This is a selected view of positive movements, not a composite score or a claim that every benchmark improved; the complete results and regressions remain in the evaluation table that follows.
Evaluation Results
Five evaluation runs. Runs R and A cover both release models, all 0-shot, with lm-evaluation-harness 0.4.13 in float32. Run C covers the base model (revision a5bb6bcd). Run B is the BFCL v4 single-turn function-calling evaluation of the instruction-tuned model, using its trained tool-calling prompt layout. Run I is IFEval for the instruction-tuned model: all 541 prompts, official lm-eval 0.4.13 scoring, the model's chat template, greedy decoding, 2,048-token response budget.
| Benchmark | Metric / shots | Base | Instruct |
|---|---|---|---|
| HellaSwag | acc_norm · 0-shot | 43.66 | 46.64 |
| ARC-Easy | acc_norm · 0-shot | 54.76 | 54.12 |
| ↳ ARC-Easy (second metric) | acc · 0-shot | 61.20 | — |
| ARC-Challenge | acc_norm · 0-shot | 29.27 | 28.41 |
| PIQA | acc_norm · 0-shot | 67.85 | 68.88 |
| WinoGrande | acc · 0-shot | 52.41 | 50.75 |
| OpenBookQA | acc_norm · 0-shot | 35.00 | 36.20 |
| BoolQ | acc · 0-shot | 49.82 | 53.36 |
| SIQA | acc · 0-shot | 40.89 | 40.33 |
| CommonsenseQA | acc · 0-shot | 19.98 | 18.18 |
| SciQ | acc · 0-shot | 82.90 | 82.70 |
| ↳ SciQ (second metric) | acc_norm · 0-shot | 75.80 | 77.10 |
| MMLU | acc · 0-shot | 25.31 | 24.78 |
| MMLU | accuracy · 5-shot | 25.36 | — |
| GPQA-Diamond | accuracy · 0-shot | 26.77 | — |
| BFCL simple (simple_python) | AST accuracy · 0-shot | — | 59.00 |
| BFCL multiple | AST accuracy · 0-shot | — | 56.00 |
| BFCL irrelevance | irrelevance (official rule) · 0-shot | — | 2.50 |
| IFEval, prompt-level strict | accuracy · 0-shot | — | 22.92 |
| IFEval, instruction-level strict | accuracy · 0-shot | — | 37.77 |
| IFEval, prompt-level loose | accuracy · 0-shot | — | 24.95 |
| IFEval, instruction-level loose | accuracy · 0-shot | — | 40.53 |
Scores are percentages. Base = the final annealed base model in this repository; Instruct = HobbyLM-1B Instruct. Task versions, example counts, the run behind each score and standard errors are in Detailed results below the table.
Detailed results: task versions, example counts, run (R, A, C, B, I), standard errors and Instruct − base differences
| Benchmark | Metric | Shots | lm-eval task version | Examples (each model) | Run | Base, % ± SE | Instruct, % ± SE | Instruct − base, percentage points |
|---|---|---|---|---|---|---|---|---|
| HellaSwag | acc_norm | 0 | 1.0 | 10,042 | R | 43.66 ± 0.49 | 46.64 ± 0.50 | +2.99 |
| ARC-Easy | acc_norm | 0 | 1.0 | 2,376 | R | 54.76 ± 1.02 | 54.12 ± 1.02 | −0.63 |
| ↳ ARC-Easy (second metric) | acc | 0 | — | — | C | 61.20 | — | — |
| ARC-Challenge | acc_norm | 0 | 1.0 | 1,172 | R | 29.27 ± 1.33 | 28.41 ± 1.32 | −0.85 |
| PIQA | acc_norm | 0 | 1.0 | 1,838 | R | 67.85 ± 1.09 | 68.88 ± 1.08 | +1.03 |
| WinoGrande | acc | 0 | 1.0 | 1,267 | R | 52.41 ± 1.40 | 50.75 ± 1.41 | −1.66 |
| OpenBookQA | acc_norm | 0 | 1.0 | 500 | R | 35.00 ± 2.14 | 36.20 ± 2.15 | +1.20 |
| BoolQ | acc | 0 | 2.0 | 3,270 | R | 49.82 ± 0.87 | 53.36 ± 0.87 | +3.55 |
| SIQA | acc | 0 | 0.0 | 1,954 | A | 40.89 ± 1.11 | 40.33 ± 1.11 | −0.56 |
| CommonsenseQA | acc | 0 | not declared | 1,221 | A | 19.98 ± 1.14 | 18.18 ± 1.10 | −1.80 |
| SciQ | acc | 0 | 1.0 | 1,000 | A | 82.90 ± 1.19 | 82.70 ± 1.20 | −0.20 |
| ↳ SciQ, same 1,000 questions (second metric) | acc_norm | 0 | 1.0 | 1,000 | A | 75.80 ± 1.36 | 77.10 ± 1.33 | +1.30 |
| MMLU (0-shot) | acc | 0 | group 2 (subtasks 1.0) | 14,042 | A | 25.31 ± 0.37 | 24.78 ± 0.36 | −0.53 |
| MMLU (5-shot) | accuracy | 5 | — | — | C | 25.36 | — | — |
| GPQA-Diamond | accuracy | 0 | custom gpqa_diamond_zeroshot_fixed |
— | C | 26.77 | — | — |
| BFCL simple (simple_python) | AST accuracy | 0 | bfcl-eval 2026.3.23 | 400 | B | — | 59.00 | — |
| BFCL multiple | AST accuracy | 0 | bfcl-eval 2026.3.23 | 200 | B | — | 56.00 | — |
| BFCL irrelevance | irrelevance (official rule) | 0 | bfcl-eval 2026.3.23 | 240 | B | — | 2.50 | — |
| IFEval, prompt-level strict | accuracy | 0 | lm-eval 0.4.13 | 541 | I | — | 22.92 | — |
| IFEval, instruction-level strict | accuracy | 0 | lm-eval 0.4.13 | 541 | I | — | 37.77 | — |
| IFEval, prompt-level loose | accuracy | 0 | lm-eval 0.4.13 | 541 | I | — | 24.95 | — |
| IFEval, instruction-level loose | accuracy | 0 | lm-eval 0.4.13 | 541 | I | — | 40.53 | — |
How to read the table
How to read the table:
- Scores are accuracy in percent. "±" is the standard error reported by lm-evaluation-harness for each model separately.
- Differences are in percentage points, computed from unrounded scores, so they may not equal the difference of the rounded values shown.
- Observed changes only: no paired significance test was run, and no significance claim is made.
- SciQ is one benchmark with two metrics,
accandacc_norm, over the same 1,000 questions. - MMLU (0-shot, Run A) is lm-eval's group accuracy, weighted by subject size across 57 subjects.
- No overall average is reported: the benchmarks use different metrics, chance levels and difficulty.
- Chance level: CommonsenseQA (both models) is near the five-choice uniform-random baseline (20%), and MMLU near the four-choice baseline (25%). No cause is attributed.
- Runs R and A used every example of each task's evaluation split.
Evaluation setup and provenance
Two historical figures are deliberately left out because no saved result artifact backs them:
- the 7-task average 47.61 shown on earlier revisions of this card;
- the function-calling elicitation pass@k numbers.
Scores of the Instruct model or of any other HobbyLM checkpoint do not apply to this base model.
Common setup (both runs):
- lm-evaluation-harness 0.4.13,
HFLMwith the model loaded viatrust_remote_code=True; - 0-shot, batch size 8, seeds 1234 (
random_seed,numpy_random_seed,torch_random_seed,fewshot_random_seed); - float32 with TF32 disabled, on one NVIDIA A100-SXM4-40GB (Modal);
- no chat template; this model's context is 1,024 tokens.
- Weights: base
harims95/hobbylm-1b-hf(nowharims95/hobbylm-1B, this repository) @a5bb6bcdab3642acbb203c2b6b6c272669da3693; instructharims95/hobbylm-1b-broad-sft-3450-hf@ddf46d8f9c651ca3d0a73bb3189a7dc9a9112ce5, the same weights asharims95/hobbylm-1B-instruct. - Tokenizer: the release tokenizer files staged for this repo.
Run R: setup and provenance limitations
Software: Python 3.11, torch 2.4.1+cu121, transformers 4.46.3, accelerate 1.1.1, safetensors 0.4.5. The
datasetsandevaluateversions were not recorded.Datasets (as configured by lm-eval 0.4.13; revisions not pinned or recorded):
Task Dataset Split HellaSwag Rowan/hellaswagvalidation ARC-Easy, ARC-Challenge allenai/ai2_arctest PIQA baber/piqavalidation WinoGrande allenai/winograndewinogrande_xlvalidation OpenBookQA allenai/openbookqamaintest BoolQ aps/super_glueboolqvalidation Tokenizer: the runs did not record tokenizer hashes. The files published here are the current staged files (SHA-256:
tokenizer.json8414cab9…,vocab.json19613966…,merges.txt1ce16647…,special_tokens_map.json259f2e7d…,tokenizer_config.jsonbase7e1ab496…/ Instruct2a199fc0…). Their timestamps predate the runs, but that alone does not prove they are byte-identical to what was loaded.Run-1 script: the evaluation script was edited after the base model's run started (the edit affected only HumanEval, which was not evaluated). Its run-time source is a reconstruction, not a captured copy.
BoolQ truncation (base model): 1 of 3,270 BoolQ documents (doc id 3154) exceeds the base model's 1,024-token context. As pre-specified, lm-eval left-truncated it, and it stays in the score. Without it, base BoolQ is 49.83% (n = 3,269). No Instruct document exceeded 4,096 tokens.
Run A: setup and verified provenance (applies to Run A only)
These checks ran inside the Run A job, before any model was loaded. They do not cover Run R.
Executed script: SHA-256
d54d2097…, verified in the container against the hash recorded before launch.Dependencies: all 90 pins of a recorded lock file verified at exact versions (torch 2.4.1, transformers 4.46.3, lm-eval 0.4.13, datasets 5.0.1, evaluate 0.4.6, huggingface-hub 0.36.2).
Tokenizer: all 10 staged tokenizer files verified against recorded SHA-256 values.
Models: both resolved to the pinned commits above.
Datasets: loaded only at pinned revisions:
Task Dataset @ revision Split SIQA lighteval/siqa@54c6a1f8cb6daf4f5abf24a601852612fb35eb25validation CommonsenseQA tau/commonsense_qa@94630fe30dad47192a8546eb75f094926d47e155validation SciQ allenai/sciq@2c94ad3e1aafab77146f384e23536f97a4849815test MMLU cais/mmlu@c30699e8356da336a370243923dbaf21066bb9fetest Context: no document exceeded either model's context, so nothing was truncated.
Run I (IFEval)
- Cap hits: 135 of 541 responses (25.0%) reached the 2,048-token limit; most were flagged as repetitive by a simple heuristic.
- Training overlap: 2 prompts overlap the final instruction-tuning mix. The scores include them.
Not evaluated: GSM8K, HumanEval, MBPP, MATH-500, TriviaQA, MMLU-Pro.
Training Recipe
The full-scale run followed an earlier proxy run that exposed a data-delivery problem: source families were reaching the model in long sequential stretches. The loader was corrected and re-tested before scaling. The illustration below is conceptual; its proportions and ordering are not literal.
- Tokens and hardware: 100B tokens, trained from scratch on 4× H200 for about 76 hours. Muon optimizer, trapezoidal learning-rate schedule.
- Mix:
| Source | Share | Dataset |
|---|---|---|
| Educational web | 60% | FineWeb-Edu |
| Web | 15% | DCLM-baseline-1.0 |
| Code | 10% | codeparrot-clean |
| Math | 10% | FineMath (finemath-4plus) |
| Anneal | 5% | Cosmopedia v2 (3.5B) + FineMath-4plus (1.5B) |
Correction: the previous card described the anneal as "Cosmopedia + high-quality FineWeb-Edu". The data-preparation script (prepare_mix100B.py) shows Cosmopedia v2 + FineMath-4plus.
Validation loss (FineWeb validation set): 3.4112 at the end of the main pretraining phase (step 81,060), and 3.5487 at the end of the full run after the anneal (step 95,367, the run's result.json). The chart below shows the main pretraining phase only.
After pretraining: the context was extended to 2048 and then 4096 tokens, a short function-calling stage followed, and broad instruction fine-tuning ran to step 3450. Every stage uses plain RoPE. The final base, the 4K context-extension checkpoint and the instruction-tuned model, with hashes and lineage, are in hobbylm-1b-checkpoints.
Quickstart
Transformers (this base model)
trust_remote_code=True is required, and the model must stay in float32 (see Limitations). This repository has no chat template.
Executed (2026-10-06, CPU, transformers 4.46.3, model from published revision a5bb6bcd; tokenizer from the STAGED files, because they are not published yet): ran without error.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "harims95/hobbylm-1B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True, torch_dtype=torch.float32).eval()
ids = tok("The three states of matter are", return_tensors="pt").input_ids
out = model.generate(ids, attention_mask=torch.ones_like(ids), max_new_tokens=40, do_sample=False, pad_token_id=tok.eos_token_id)
print(tok.decode(out[0]))
Pinned revisions and tokenizer check
Revisions.
- Quick start: load without
revision=(the latest revision has the weights, model code, tokenizer and this card). - Exact reproduction needs immutable revisions for both parts:
- model: weights revision
a5bb6bcdab3642acbb203c2b6b6c272669da3693, the revision the evaluations in this card used; - tokenizer: that revision has no tokenizer files. They were added in a later commit that did not change the weights,
config.jsonor the model code. Load the tokenizer from that commit,dc11aab4e06fd2000313c8821557a86e8368bf58, and check the files against these SHA-256 values:merges.txt1ce1664773c50f3e0cc8842619a93edc4624525b728b188a9e0be33b7726adc5;special_tokens_map.json259f2e7dd02184bad76951125cc8b79cd2f3a8c44ca7dc66737742e5e4f8ed9b;tokenizer.json8414cab924d8b9b33013f0d221c5862f365ee9be39c5c2bfae8a5a9e970478a6;tokenizer_config.json7e1ab4968c5c4e97959fb1f6fe5bd979cece96d1a99ac1f15e50c95116585a23;vocab.json196139668be63f3b5d6574427317ae82f612a97c5d1cdaf36ed2256dbf636783.
- model: weights revision
- Example:
AutoModelForCausalLM.from_pretrained(repo, revision="a5bb6bcdab3642acbb203c2b6b6c272669da3693", trust_remote_code=True, torch_dtype=torch.float32)andAutoTokenizer.from_pretrained(repo, revision="dc11aab4e06fd2000313c8821557a86e8368bf58").
Tested (CPU, tokenizer only, transformers 4.46.3): the bundled tokenizer encodes and decodes token-for-token identically to tiktoken gpt2, and adds no BOS/EOS.
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("harims95/hobbylm-1B")
assert tok.encode("Hello") == [15496]
Instruct model
The instruction-tuned model has its own repository and card: HobbyLM-1B Instruct.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo, rev = "harims95/hobbylm-1B-instruct", "0506eed260c353a259712705fa2eb11662f1f0ac"
tok = AutoTokenizer.from_pretrained(repo, revision=rev)
model = AutoModelForCausalLM.from_pretrained(repo, revision=rev, trust_remote_code=True, torch_dtype=torch.float32)
text = tok.apply_chat_template([{"role": "user", "content": "What is the capital of France?"}], tokenize=False, add_generation_prompt=True)
ids = tok(text, return_tensors="pt")
out = model.generate(**ids, max_new_tokens=64, do_sample=False, pad_token_id=50256)
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True)) # The capital of France is Paris.
GGUF with the patched llama.cpp
The GGUF files need the patched runtime from the v1.0.0 release; stock llama.cpp and stock Ollama cannot load them.
huggingface-cli download harims95/hobbylm-1B-gguf hobbylm-1b-sft-3450-Q4_K_M.gguf --revision f8bc034403026b2706f5275702d3801753b87c34 --local-dir .
start-chat-server.bat "C:\path\to\hobbylm-1b-sft-3450-Q4_K_M.gguf"
Then open http://127.0.0.1:8080/ in your browser. The base model runs with run-base-completion.bat. Ollama steps, checksums and test results: the GGUF model card.
Older download commands: commands that pin harims95/hobbylm-1b-hf (the former name of this repository, which redirects here) at revision dc11aab4e06fd2000313c8821557a86e8368bf58 keep working, because that revision still contains the GGUF files.
Windows chat app
Download HobbyLM-1B-Chat-Windows-x64-1.0.0.zip from the release page, extract it, and double-click Start HobbyLM.bat. The chat opens in your browser and runs offline on your computer.
About the Windows packages
All three packages are unofficial builds for Windows x86-64, CPU only. They are not produced, reviewed or endorsed by the llama.cpp/ggml or Ollama projects, and their rebuilt binaries are not code-signed.
| Package | What it is | Tested |
|---|---|---|
| HobbyLM-1B Chat for Windows (double-click bundle) | A ZIP containing Start HobbyLM.bat, the Instruct Q4_K_M GGUF and the standalone llama.cpp runtime below. It chats in the browser on 127.0.0.1 only. |
Extract, launch, chat, close and relaunch; the same-folder launch guard. |
| Standalone patched llama.cpp runtime | llama.cpp 63bef272 plus 2 HobbyLM patches: the QKV-shape fix and a Windows exit fix. Includes llama-server with a web UI built from source, llama-completion and launchers. No model included. |
Recorded outputs reproduced from a clean PATH; offline; web UI checked in a browser. |
| Patched Ollama runtime | The official ollama.exe v0.35.1, unchanged, plus a runtime rebuilt from Ollama's llama.cpp (b11232) with 2 HobbyLM patches. Uses an isolated address (127.0.0.1:11435) and data folder. No model included. |
Recorded outputs reproduced through the raw and chat APIs. |
- Stock runtimes: stock Ollama v0.35.1 was tested and fails to load the model (
attn_qkv.weightshape mismatch). Unpatched llama.cpp was not run; its code expects the same QKV shape that Ollama rejects. - Test coverage: everything above ran on one machine: Windows 11, AMD Ryzen 9 5900HX, CPU backend, AVX2 ("haswell") code path.
- Not tested: the other x86-64 CPU variants included in the packages, Windows on ARM, Linux, macOS, and any GPU backend (CUDA, Vulkan, Metal).
- Downloads: HobbyLM 1.0.0 release on GitHub: the double-click chat bundle, the patched llama.cpp and Ollama runtimes, and their source archive.
Limitations and Disclaimer
- Intended use: research on small sparse MoE models, pretraining and fine-tuning experiments. Not for: production use or as a factual reference.
- No safety tuning or preference tuning. Outputs can be wrong, biased or harmful.
- Context: the base model was trained at 1024 tokens. The instruction-tuned model is configured for 4096 tokens, but its 4K parent failed a 4K retrieval certification, so long-context recall is not demonstrated.
- Precision: The FP32 Transformers models (the root files of this repository,
hobbylm-1B-instructand the checkpoints inhobbylm-1b-checkpoints) must stay fp32. Converting the router to bf16 changes which experts are selected: router top-1 agreement drops to 62%, and full-model output agreement drops from 97.6% to 61%. With top-8-of-64 routing, the score gap between rank 8 and rank 9 is below bf16's precision, so lowering the router's precision silently changes behavior. GGUF conversions (hobbylm-1B-gguf) are separate files: three formats of the instruct model (F32, Q8_0, Q4_K_M) and a separate base-model F32 file. The quantised GGUFs retain the router and expert-bias tensors in F32. Output agreement, token-identical to the FP32 Transformers reference on seven fixed prompts: Instruct F32 6/7; Instruct Q8_0 7/7; Instruct Q4_K_M 6/7. Perplexity change relative to the Instruct F32 GGUF, measured on the model's own evaluation prompts and outputs (so it shows relative quantisation loss only): +0.02% (Q8_0) and +1.1% (Q4_K_M). No lm-eval benchmarks were run on the GGUF versions, and these checks do not establish general quality parity. Details: the GGUF model card. - Tokenizer: GPT-2 BPE, English-centric and less token-efficient.
- Runtimes: vLLM and SGLang have not been validated. GGUF runs only with the supplied patched runtimes (see the compatibility note under Model List), and only the checks listed on the GGUF model card were run.
License
Apache License 2.0 (see LICENSE). Attributions are in NOTICE.
| Artifact | Licence |
|---|---|
Model weights: the base model.safetensors in this repo, and the GGUF files in hobbylm-1B-gguf |
Apache-2.0 |
Model code in this repo (configuration_hobbylm.py, modeling_hobbylm.py): a Hugging Face port of the HobbyLM architecture and original model code by Harish (github.com/harishsg993010/HobbyLM) |
Apache-2.0, released with its author's permission |
Tokenizer vocabulary files (vocab.json, merges.txt, tokenizer.json): GPT-2 BPE from openai-community/gpt2 |
MIT (OpenAI), unchanged |
| Training data | each source's own terms (see NOTICE and below) |
| Runtime software (e.g. llama.cpp, Ollama) used to run the GGUF files | its own licences; not covered by this licence |
Training-data licence notes
Apache-2.0 is the authors' licence for the weights. It does not change the terms of the training data, and it does not resolve this open question, disclosed here:
- codeparrot-clean (10% of pretraining) declares no dataset licence; each file carries its original repository's licence. No licence filter was applied, so copyleft-licensed files may be included.
Other pretraining sources (FineWeb-Edu, FineMath, Cosmopedia v2: ODC-By 1.0; DCLM-baseline: CC BY 4.0) are attributed in NOTICE.
Credits
Built by Hariharan and Prabhurajhan at Fuel Labs. Architecture based on HobbyLM by Harish.
- Downloads last month
- 61


