Instructions to use startlux-models/Lodestar-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use startlux-models/Lodestar-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="startlux-models/Lodestar-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("startlux-models/Lodestar-4B") model = AutoModelForMultimodalLM.from_pretrained("startlux-models/Lodestar-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use startlux-models/Lodestar-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "startlux-models/Lodestar-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "startlux-models/Lodestar-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/startlux-models/Lodestar-4B
- SGLang
How to use startlux-models/Lodestar-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "startlux-models/Lodestar-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "startlux-models/Lodestar-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "startlux-models/Lodestar-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "startlux-models/Lodestar-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use startlux-models/Lodestar-4B with Docker Model Runner:
docker model run hf.co/startlux-models/Lodestar-4B
Lodestar-4B
Lodestar-4B is a 4-billion-parameter decision model. You give it a state (plain text, JSON or a long document) and one or more typed questions; it returns a probability for every listed option. Each answer is read from a single forward pass. The model never generates free text, so there is nothing to parse and no output length to budget for.
| Base model | Qwen3.5-4B-Base (32 layers: 24 Gated DeltaNet, 8 full attention) |
| Question types | choice (one of up to 26 options per pass; longer lists are resolved in rounds), noul (yes/no), score (ordinal levels) |
| Output | a probability for every option, calibrated per question type |
| Interface | Python API and a TypeSafe-compatible POST /v1/systemone server (included) |
| Latency | 11 ms per short question, 21 ms median on the multi-thousand-token hard items (one H200, bf16, CUDA graphs) |
| Context | tested up to 64k tokens |
| Precision | bf16, about 9 GB of weights |
| License | Apache-2.0 |
Quick start
hf download startlux-models/Lodestar-4B --local-dir Lodestar-4B
cd Lodestar-4B
pip install -r requirements.txt # flash-linear-attention is optional but much faster
python -m lodestar.server --model . --port 8090
curl -s localhost:8090/v1/systemone -H 'Content-Type: application/json' -d '{
"state": {"ticket": "I was charged twice for order #4411 and the app still shows it as unpaid."},
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Payments, refunds and invoices",
"shipping": "Delivery and tracking",
"technical": "App, login and account problems"}},
"urgent": {"type": "noul", "instructions": "Should this ticket be answered today?"}
}}'
{"answers": {"team": {"type": "choice", "choice": "billing", "probabilities": {"billing": 0.965, "shipping": 0.002, "technical": 0.034}},
"urgent": {"type": "noul", "noul": 0.623}},
"usage": {"input_tokens": 191, "output_tokens": 0}, "model": "Lodestar-4B"}
(Probabilities rounded to three decimals.) The same call from Python, run inside the downloaded folder:
from lodestar import Lodestar
model = Lodestar(".") # or "startlux-models/Lodestar-4B"; needs one CUDA GPU
answers, usage = model.decide(state, questions)
A score question takes its levels as a list, lowest first, and returns
{"type": "score", "score": <level index>, "probabilities": {"0": p0, "1": p1, ...}}.
How it answers
Every question is rendered into one chat prompt (thinking disabled):
system: Apply the criterion to the evidence. Choose exactly one listed option. Answer with its letter only.
user: Evidence:
<state>
Question: <instructions>
Options:
A) <option id>: <description>
B) ...
The probability of each option is the softmax of the next-token logits of its letter, taken at the last prompt
position and divided by the temperature of the question type (lodestar_config.json). Yes/no questions are shown as
the options yes and no. Choice lists longer than 26 options are split into groups of 25; the top three of each group
go to a final round, and the options that miss it share a small residual probability. The renderer is
lodestar/jevfmt.py; prompts built another way will not reproduce the published behaviour.
Training
- Supervised fine-tuning, all parameters. About 2.2 million typed decisions, converted into the prompt format above from publicly available datasets: intent and topic classification, natural-language inference, reading comprehension and multiple-choice knowledge questions, policy and rule application, tool and routing selection, answer verification, preference and quality judgments, field extraction, and numeric and temporal reasoning. Where a source provides a label distribution rather than a single label, the model is trained toward that distribution. One epoch, learning rate 2e-5 with cosine decay.
- Targeted refinement. A rank-64 LoRA over the attention, DeltaNet and MLP projections, trained for two epochs on 61k rows: decisions the first-stage model got wrong or was unsure about, human-labelled sets, and replay rows whose target is the first-stage model's own distribution, so that the refinement does not erode what the model already does well. The adapter is merged into the released weights.
- Calibration. One temperature per question type, fitted on held-out hard items:
noul1.79,choice1.47,score1.51. Temperatures change confidence, never the chosen option.
Training data was decontaminated against the benchmark items used for evaluation (13-gram overlap and exact match): the Decision Index 0.2 test splits, the JevBench public set and Humanity's Last Exam. A separate exhaustive 13-gram check of both training stages against the 231 JevBench public items finds no overlapping example.
Evaluation
JevBench public split (231 items), run through JevBench's own typesafe adapter against the bundled server:
| Tier | Items | Correct |
|---|---|---|
| Easy | 48 | 48 |
| Standard | 72 | 70 |
| Hard | 111 | 80 |
| All | 231 | 198 |
These are self-reported results on the public items. They are not an official leaderboard entry.
Limitations
- One pass, no chain of thought: multi-step arithmetic, date calculations and subtle answer checking are the weakest areas, especially inside long documents.
- The temperatures were fitted on hard items, so on easy inputs the probabilities are on the cautious side.
- Training data is predominantly English.
- The model decides among the options it is given. It does not add options, and it answers
noulquestions even when the evidence is insufficient; give it an explicit "cannot tell" option when that matters. - It is a decision component, not a chat assistant.
License
The weights are released under Apache-2.0, following the base model. The training data comes from public datasets, each under its own license.
- Downloads last month
- -
Model tree for startlux-models/Lodestar-4B
Base model
Qwen/Qwen3.5-4B-Base