MLX
Safetensors
lora
calibration
typed-decisions
apple-silicon

LARS: Limited but Accurate Response System

Weights for LARS, a local engine for typed decisions that speaks the System One API. You send a state (text or JSON) and typed questions: a yes/no claim, a choice among options, an ordered score, or a multi-label pick. LARS answers each one with calibrated probabilities read from the model's label logits. It generates no text. It runs on Apple Silicon with MLX.

Two tiers live here, each in its own folder with its pooled calibrator:

folder what it is preset resident memory
4b/ Qwen3-4B-Instruct-2507 4-bit with its LoRA fused in (the adapter's layers at 8 bits, the rest at 4) lars serve --preset 4b about 5 GB
30b-a3b/ an attention-only LoRA for mlx-community/Qwen3-30B-A3B-Instruct-2507-4bit lars serve --preset 30b about 18 GB

Each folder has a lars-calibration.json: temperature and Platt scaling per question type, fitted on pooled validation rows from all 22 benchmark configs. The presets load it automatically.

Use

uv tool install "lars-engine[mlx,mcp] @ git+https://github.com/moedex/lars"
lars serve --preset 30b                  # or --preset 4b on smaller machines
curl -s localhost:8600/v1/systemone -H 'content-type: application/json' -d '{
  "state": "The build failed: 3 tests in test_auth.py timed out after the Redis upgrade.",
  "questions": {"flaky": {"type": "noul", "instructions": "The failure is caused by a flaky test"}}
}'

LARS also serves MCP at /mcp, and lars service install runs it at login. See the repository for the API, per-decision calibrators, and optional escalation to a hosted model for questions the local model doesn't know.

Evaluation

jev-bench, 22 configs, 200 test rows each, served exactly as the presets serve it (one pooled calibrator for every config). The macro average weights each config equally. Jev's numbers are its published ones.

system macro accuracy Brier ECE
LARS 30B-A3B 0.767 0.285 0.082
LARS 4B 0.741 0.305 0.074
Jev (published) 0.733 0.349 0.113

Per config (calibrated accuracy, as served):

config type 30B-A3B 4B Jev
banking77 choice 0.820 0.835 0.796
boolq noul 0.895 0.900 0.917
sst5 (unseen) score 0.500 0.385 0.565
clinc150 choice 0.915 0.920 0.893
massive choice 0.890 0.865 0.808
ledgar choice 0.805 0.810 0.751
go_emotions choice 0.515 0.525 0.282
mmlu choice 0.785 0.700 0.923
arc_challenge choice 0.950 0.915 0.979
mnli choice 0.850 0.860 0.883
chaosnli (unseen) choice 0.645 0.650 0.615
yelp5 (unseen) score 0.600 0.565 0.685
helpsteer2_helpfulness (unseen) score 0.370 0.285 0.363
helpsteer2_verbosity score 0.680 0.675 0.341
stsb (unseen) score 0.435 0.285 0.538
measuring_hate_speech score 0.775 0.805 0.527
fever_evidence (unseen) noul 0.950 0.925 0.972
paws noul 0.915 0.915 0.846
civil_comments (unseen) noul 0.925 0.930 0.729
sms_spam noul 0.965 0.960 0.965
strategyqa_closed noul 0.755 0.680 0.785
strategyqa_grounded noul 0.925 0.915 0.956

How to read this:

  • Unseen configs: civil_comments, fever_evidence and helpsteer2_helpfulness were held out of training entirely. sst5, yelp5 and stsb were dropped from training for license reasons (below), and chaosnli has no training source. These seven measure generalization. The rest share sources with training rows, though never the test rows.
  • Where LARS wins and loses: the wins are large on classification and moderation (go_emotions, helpsteer2_verbosity, measuring_hate_speech, civil_comments, massive). LARS is behind Jev on knowledge-heavy tasks (mmlu, strategyqa_closed, arc_challenge) and fine-grained scores (sst5, stsb). The 4B is further behind there: a 4B model doesn't hold that knowledge.
  • Sample size: 200 test rows per config gives intervals of several points on any one config. The macro is steadier. Paired bootstrap intervals are in the repository's evals/RESULTS.md.

Training

Both tiers were trained on the same corpus: 37,962 rows, one epoch, lr 2e-5, rank 8, scale 20. The loss is cross-entropy plus Brier on the label logits at the answer position. The 4B adapter covers every linear layer; the 30B adapter covers attention only. The checkpoint was chosen on validation rows, weighting seen and held-out sources equally. Seed 0 for both.

source rows license
ZefanCai/Open-Jev, release-v2-redistributable 18,777 CC0 1.0
jev-bench train splits 15,185 per source, below
tasksource-jev: SNLI, MultiNLI, bAbI NLI 4,000 CC BY-SA 4.0, mixed OANC / CC BY-SA 3.0, BSD

jev-bench training sources and their licenses:

  • banking77, massive, ledgar, helpsteer2_verbosity, measuring_hate_speech and sms_spam are CC BY 4.0.
  • clinc150 is CC BY 3.0.
  • arc_challenge is CC BY-SA 4.0, and boolq is CC BY-SA 3.0.
  • mnli is MultiNLI: mostly OANC, with some CC BY-SA 3.0 and CC BY 3.0.
  • go_emotions is Apache 2.0.
  • paws is free for any purpose (Google).
  • strategyqa and mmlu are MIT.

Excluded from training:

  • yelp5, which is non-commercial and not redistributable.
  • sst5 and stsb, whose licenses are unverified.
  • ANLI, which is CC BY-NC.
  • the held-out configs above.
  • every eval row.

Data policy:

  • No training row carries outputs from Jev or any other hosted model. The labels are human gold labels and vote shares from the original datasets, plus Open-Jev's synthetic controls.
  • Whether share-alike terms reach trained weights is untested. The weights are published under Apache 2.0, with every source attributed above.

The base models are Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507 (Apache 2.0, Qwen team), via their mlx-community 4-bit conversions.

Limitations

  • Typed decisions only. It won't write, summarize or explain. Answers are probabilities over the options you give.
  • Knowledge is limited by model size. Questions that need facts the model doesn't hold come back less confident, and calibration keeps that visible. serve --escalate can send those rows to a hosted model when you choose to allow it.
  • Your own decisions need your own validation rows. The pooled calibrator is fitted on benchmark configs. For a decision of your own, fit a calibrator on your labeled rows (lars calibrate) and check accuracy before trusting its probabilities.
  • Apple Silicon only. These folders are MLX weights. The repository's llama.cpp backend runs GGUF models elsewhere, but GGUF versions of these tiers aren't published yet.
Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for moedex/lars

Adapter
(128)
this model

Datasets used to train moedex/lars