Instructions to use moedex/lars with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use moedex/lars with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir lars moedex/lars
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
LARS: Limited but Accurate Response System
Weights for LARS, a local engine for typed decisions that speaks the System One API. You send a state (text or JSON) and typed questions: a yes/no claim, a choice among options, an ordered score, or a multi-label pick. LARS answers each one with calibrated probabilities read from the model's label logits. It generates no text. It runs on Apple Silicon with MLX.
Two tiers live here, each in its own folder with its pooled calibrator:
| folder | what it is | preset | resident memory |
|---|---|---|---|
4b/ |
Qwen3-4B-Instruct-2507 4-bit with its LoRA fused in (the adapter's layers at 8 bits, the rest at 4) | lars serve --preset 4b |
about 5 GB |
30b-a3b/ |
an attention-only LoRA for mlx-community/Qwen3-30B-A3B-Instruct-2507-4bit |
lars serve --preset 30b |
about 18 GB |
Each folder has a lars-calibration.json: temperature and Platt scaling per question type,
fitted on pooled validation rows from all 22 benchmark configs. The presets load it
automatically.
Use
uv tool install "lars-engine[mlx,mcp] @ git+https://github.com/moedex/lars"
lars serve --preset 30b # or --preset 4b on smaller machines
curl -s localhost:8600/v1/systemone -H 'content-type: application/json' -d '{
"state": "The build failed: 3 tests in test_auth.py timed out after the Redis upgrade.",
"questions": {"flaky": {"type": "noul", "instructions": "The failure is caused by a flaky test"}}
}'
LARS also serves MCP at /mcp, and lars service install runs it at login. See the
repository for the API, per-decision calibrators, and
optional escalation to a hosted model for questions the local model doesn't know.
Evaluation
jev-bench, 22 configs, 200 test rows each, served exactly as the presets serve it (one pooled calibrator for every config). The macro average weights each config equally. Jev's numbers are its published ones.
| system | macro accuracy | Brier | ECE |
|---|---|---|---|
| LARS 30B-A3B | 0.767 | 0.285 | 0.082 |
| LARS 4B | 0.741 | 0.305 | 0.074 |
| Jev (published) | 0.733 | 0.349 | 0.113 |
Per config (calibrated accuracy, as served):
| config | type | 30B-A3B | 4B | Jev |
|---|---|---|---|---|
| banking77 | choice | 0.820 | 0.835 | 0.796 |
| boolq | noul | 0.895 | 0.900 | 0.917 |
| sst5 (unseen) | score | 0.500 | 0.385 | 0.565 |
| clinc150 | choice | 0.915 | 0.920 | 0.893 |
| massive | choice | 0.890 | 0.865 | 0.808 |
| ledgar | choice | 0.805 | 0.810 | 0.751 |
| go_emotions | choice | 0.515 | 0.525 | 0.282 |
| mmlu | choice | 0.785 | 0.700 | 0.923 |
| arc_challenge | choice | 0.950 | 0.915 | 0.979 |
| mnli | choice | 0.850 | 0.860 | 0.883 |
| chaosnli (unseen) | choice | 0.645 | 0.650 | 0.615 |
| yelp5 (unseen) | score | 0.600 | 0.565 | 0.685 |
| helpsteer2_helpfulness (unseen) | score | 0.370 | 0.285 | 0.363 |
| helpsteer2_verbosity | score | 0.680 | 0.675 | 0.341 |
| stsb (unseen) | score | 0.435 | 0.285 | 0.538 |
| measuring_hate_speech | score | 0.775 | 0.805 | 0.527 |
| fever_evidence (unseen) | noul | 0.950 | 0.925 | 0.972 |
| paws | noul | 0.915 | 0.915 | 0.846 |
| civil_comments (unseen) | noul | 0.925 | 0.930 | 0.729 |
| sms_spam | noul | 0.965 | 0.960 | 0.965 |
| strategyqa_closed | noul | 0.755 | 0.680 | 0.785 |
| strategyqa_grounded | noul | 0.925 | 0.915 | 0.956 |
How to read this:
- Unseen configs: civil_comments, fever_evidence and helpsteer2_helpfulness were held out of training entirely. sst5, yelp5 and stsb were dropped from training for license reasons (below), and chaosnli has no training source. These seven measure generalization. The rest share sources with training rows, though never the test rows.
- Where LARS wins and loses: the wins are large on classification and moderation (go_emotions, helpsteer2_verbosity, measuring_hate_speech, civil_comments, massive). LARS is behind Jev on knowledge-heavy tasks (mmlu, strategyqa_closed, arc_challenge) and fine-grained scores (sst5, stsb). The 4B is further behind there: a 4B model doesn't hold that knowledge.
- Sample size: 200 test rows per config gives intervals of several points on any one
config. The macro is steadier. Paired bootstrap intervals are in the repository's
evals/RESULTS.md.
Training
Both tiers were trained on the same corpus: 37,962 rows, one epoch, lr 2e-5, rank 8, scale 20. The loss is cross-entropy plus Brier on the label logits at the answer position. The 4B adapter covers every linear layer; the 30B adapter covers attention only. The checkpoint was chosen on validation rows, weighting seen and held-out sources equally. Seed 0 for both.
| source | rows | license |
|---|---|---|
ZefanCai/Open-Jev, release-v2-redistributable |
18,777 | CC0 1.0 |
| jev-bench train splits | 15,185 | per source, below |
| tasksource-jev: SNLI, MultiNLI, bAbI NLI | 4,000 | CC BY-SA 4.0, mixed OANC / CC BY-SA 3.0, BSD |
jev-bench training sources and their licenses:
- banking77, massive, ledgar, helpsteer2_verbosity, measuring_hate_speech and sms_spam are CC BY 4.0.
- clinc150 is CC BY 3.0.
- arc_challenge is CC BY-SA 4.0, and boolq is CC BY-SA 3.0.
- mnli is MultiNLI: mostly OANC, with some CC BY-SA 3.0 and CC BY 3.0.
- go_emotions is Apache 2.0.
- paws is free for any purpose (Google).
- strategyqa and mmlu are MIT.
Excluded from training:
- yelp5, which is non-commercial and not redistributable.
- sst5 and stsb, whose licenses are unverified.
- ANLI, which is CC BY-NC.
- the held-out configs above.
- every eval row.
Data policy:
- No training row carries outputs from Jev or any other hosted model. The labels are human gold labels and vote shares from the original datasets, plus Open-Jev's synthetic controls.
- Whether share-alike terms reach trained weights is untested. The weights are published under Apache 2.0, with every source attributed above.
The base models are Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507 (Apache 2.0, Qwen team), via their mlx-community 4-bit conversions.
Limitations
- Typed decisions only. It won't write, summarize or explain. Answers are probabilities over the options you give.
- Knowledge is limited by model size. Questions that need facts the model doesn't hold
come back less confident, and calibration keeps that visible.
serve --escalatecan send those rows to a hosted model when you choose to allow it. - Your own decisions need your own validation rows. The pooled calibrator is fitted on
benchmark configs. For a decision of your own, fit a calibrator on your labeled rows
(
lars calibrate) and check accuracy before trusting its probabilities. - Apple Silicon only. These folders are MLX weights. The repository's llama.cpp backend runs GGUF models elsewhere, but GGUF versions of these tiers aren't published yet.
Quantized
Model tree for moedex/lars
Base model
Qwen/Qwen3-30B-A3B-Instruct-2507