Instructions to use harims95/hobbylm-1B-instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use harims95/hobbylm-1B-instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="harims95/hobbylm-1B-instruct", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("harims95/hobbylm-1B-instruct", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use harims95/hobbylm-1B-instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "harims95/hobbylm-1B-instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harims95/hobbylm-1B-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/harims95/hobbylm-1B-instruct
- SGLang
How to use harims95/hobbylm-1B-instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "harims95/hobbylm-1B-instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harims95/hobbylm-1B-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "harims95/hobbylm-1B-instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harims95/hobbylm-1B-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use harims95/hobbylm-1B-instruct with Docker Model Runner:
docker model run hf.co/harims95/hobbylm-1B-instruct
HobbyLM-1B Instruct
The instruction-tuned HobbyLM-1B: a 1.037B-parameter sparse mixture-of-experts model with approximately 305M parameters active per token.
HobbyLM by Fuel Labs
Online Demo | Base Model | GGUF | Checkpoints | Windows App | GitHub
Overview
This repository holds the released instruction-tuned HobbyLM-1B at its root, so it loads directly without a subfolder. It is byte-identical to the sft-step3450 subfolder of hobbylm-1b-checkpoints, and it is the model behind the online demo.
- What it does: follows simple instructions, answers short questions, rewrites and extracts text, and makes single function calls in a Python-list format.
- Lineage: final annealed base (
hobbylm-1B) → 2K and 4K context extension → short function-calling stage → instruction tuning (3,450 steps). - Context: configured and trained at 4,096 tokens. Reliable retrieval across that window is not demonstrated (see Limitations).
- Precision: float32 only. Converting the router to bf16 changes which experts are selected.
Quickstart
trust_remote_code=True is required. Tested with transformers 4.46.3 and torch 2.4.1 (CPU); the model code requires 4.44.0 <= transformers < 4.47.0.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "harims95/hobbylm-1B-instruct"
rev = "0506eed260c353a259712705fa2eb11662f1f0ac" # verified revision
tok = AutoTokenizer.from_pretrained(repo, revision=rev)
model = AutoModelForCausalLM.from_pretrained(repo, revision=rev, trust_remote_code=True, torch_dtype=torch.float32).eval()
messages = [{"role": "user", "content": "What is the capital of France?"}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
ids = tok(text, return_tensors="pt")
out = model.generate(**ids, max_new_tokens=64, do_sample=False, pad_token_id=50256)
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))
# The capital of France is Paris.
The output shown is the recorded greedy output at this revision.
Prompt format
The chat template in tokenizer_config.json is plain text. Roles are system, user, assistant and tool. Each message starts with <|role|> on its own line; messages are separated by a newline; the generation prompt is <|assistant|>. Generation ends at <|endoftext|> (50256).
<|user|>
What is the capital of France?
<|assistant|>
Generation defaults (generation_config.json): greedy decoding (do_sample=false), up to 256 new tokens.
More recorded examples (greedy)
User: Extract the name, department, and years of experience from this sentence as JSON:
"Priya Nair, a senior engineer in the Infrastructure department, has 8 years of experience."
HobbyLM: {"name": "Priya Nair", "department": "Infrastructure", "years": 8}
User: Rewrite this sentence to remove the redundancy: "At this point in time, we are currently
reviewing the proposal document that was submitted to us."
HobbyLM: We are currently reviewing the proposal document that was submitted to us.
These are single recorded outputs, not averages; overall results are below.
Evaluation (this instruction-tuned model)
All scores are for this instruction-tuned model (HobbyLM-1B Instruct), in percent. The base model's scores are on the base model card, which also shows both models side by side with standard errors and the full evaluation setup.
| Benchmark | Metric · shots | Instruct |
|---|---|---|
| HellaSwag | acc_norm · 0-shot | 46.64 |
| ARC-Easy | acc_norm · 0-shot | 54.12 |
| ARC-Challenge | acc_norm · 0-shot | 28.41 |
| PIQA | acc_norm · 0-shot | 68.88 |
| WinoGrande | acc · 0-shot | 50.75 |
| OpenBookQA | acc_norm · 0-shot | 36.20 |
| BoolQ | acc · 0-shot | 53.36 |
| SIQA | acc · 0-shot | 40.33 |
| CommonsenseQA | acc · 0-shot | 18.18 |
| SciQ | acc / acc_norm · 0-shot | 82.70 / 77.10 |
| MMLU | acc · 0-shot | 24.78 |
| BFCL simple (simple_python) | AST accuracy · 0-shot | 59.00 |
| BFCL multiple | AST accuracy · 0-shot | 56.00 |
| BFCL irrelevance | irrelevance (official rule) · 0-shot | 2.50 |
| IFEval, prompt-level strict / loose | accuracy · 0-shot | 22.92 / 24.95 |
| IFEval, instruction-level strict / loose | accuracy · 0-shot | 37.77 / 40.53 |
- Setup: lm-evaluation-harness 0.4.13, float32, all examples of each split. BFCL v4 single-turn with the model's trained tool-calling prompt layout. IFEval: all 541 prompts, official scoring, this chat template, greedy, 2,048-token response budget.
- Weights evaluated: the same weights as this repository (published at the time as
harims95/hobbylm-1b-broad-sft-3450-hf@ddf46d8f). - Near chance: CommonsenseQA (20% baseline) and MMLU (25% baseline).
- BFCL irrelevance 2.50%: in BFCL's prompt layout, the model almost always calls a tool, even when none fits the request.
- IFEval: 135 of 541 responses reached the 2,048-token cap, mostly through repetition; 2 prompts overlap the final instruction-tuning mix and are included.
Limitations
- Small model: answers are often wrong or repetitive, and multi-turn chat is weak. Not a factual reference.
- No safety tuning and no preference tuning. Outputs can be wrong, biased or harmful.
- Function calling: single calls in the trained format only; weak at declining when no tool fits.
- Context: 4,096 tokens is the configured and trained length. Its 4K parent failed a 4K retrieval certification, so long-context recall is not demonstrated.
- Precision: float32 only. bf16 lowers router top-1 agreement to 62% and full-model output agreement from 97.6% to 61%.
- Tokenizer: GPT-2 BPE, English-centric.
- Runtimes: Hugging Face's automatic "Use this model" snippets, vLLM and SGLang are not validated. For CPU use, see the GGUF files, which need the patched runtimes from the v1.0.0 release.
Files
model.safetensors (float32), config.json, generation_config.json, configuration_hobbylm.py, modeling_hobbylm.py, tokenizer.json, vocab.json, merges.txt, special_tokens_map.json, tokenizer_config.json (with the chat template), LICENSE, NOTICE. SHA-256 of every file: SHA256SUMS.txt.
License
Apache License 2.0 (LICENSE) for the weights and the model code, a Hugging Face port of the HobbyLM architecture and original model code by Harish (github.com/harishsg993010/HobbyLM), released with its author's permission. The GPT-2 tokenizer files are MIT (OpenAI). Training data keeps its own terms; attributions and the open training-data licence question are in NOTICE and the base model card.
Credits
Built by Hariharan and Prabhurajhan at Fuel Labs. Architecture based on HobbyLM by Harish.
- Downloads last month
- 16