--- license: apache-2.0 library_name: mlx base_model: - Qwen/Qwen3-4B-Instruct-2507 - Qwen/Qwen3-30B-A3B-Instruct-2507 datasets: - Praveenrajus/jev-bench - ZefanCai/Open-Jev - tasksource/tasksource-jev tags: - mlx - lora - calibration - typed-decisions - apple-silicon --- # LARS: Limited but Accurate Response System Weights for [LARS](https://github.com/moedex/lars), a local engine for typed decisions that speaks the System One API. You send a state (text or JSON) and typed questions: a yes/no claim, a choice among options, an ordered score, or a multi-label pick. LARS answers each one with calibrated probabilities read from the model's label logits. It generates no text. It runs on Apple Silicon with MLX. Two tiers live here, each in its own folder with its pooled calibrator: | folder | what it is | preset | resident memory | |---|---|---|---| | `4b/` | Qwen3-4B-Instruct-2507 4-bit with its LoRA fused in (the adapter's layers at 8 bits, the rest at 4) | `lars serve --preset 4b` | about 5 GB | | `30b-a3b/` | an attention-only LoRA for `mlx-community/Qwen3-30B-A3B-Instruct-2507-4bit` | `lars serve --preset 30b` | about 18 GB | Each folder has a `lars-calibration.json`: temperature and Platt scaling per question type, fitted on pooled validation rows from all 22 benchmark configs. The presets load it automatically. ## Use ```bash uv tool install "lars-engine[mlx,mcp] @ git+https://github.com/moedex/lars" lars serve --preset 30b # or --preset 4b on smaller machines ``` ```bash curl -s localhost:8600/v1/systemone -H 'content-type: application/json' -d '{ "state": "The build failed: 3 tests in test_auth.py timed out after the Redis upgrade.", "questions": {"flaky": {"type": "noul", "instructions": "The failure is caused by a flaky test"}} }' ``` LARS also serves MCP at `/mcp`, and `lars service install` runs it at login. See the [repository](https://github.com/moedex/lars) for the API, per-decision calibrators, and optional escalation to a hosted model for questions the local model doesn't know. ## Evaluation [jev-bench](https://huggingface.co/datasets/Praveenrajus/jev-bench), 22 configs, 200 test rows each, served exactly as the presets serve it (one pooled calibrator for every config). The macro average weights each config equally. Jev's numbers are its published ones. | system | macro accuracy | Brier | ECE | |---|---|---|---| | **LARS 30B-A3B** | **0.767** | **0.285** | 0.082 | | **LARS 4B** | **0.741** | **0.305** | **0.074** | | Jev (published) | 0.733 | 0.349 | 0.113 | Per config (calibrated accuracy, as served): | config | type | 30B-A3B | 4B | Jev | |---|---|---|---|---| | banking77 | choice | 0.820 | 0.835 | 0.796 | | boolq | noul | 0.895 | 0.900 | 0.917 | | sst5 (unseen) | score | 0.500 | 0.385 | 0.565 | | clinc150 | choice | 0.915 | 0.920 | 0.893 | | massive | choice | 0.890 | 0.865 | 0.808 | | ledgar | choice | 0.805 | 0.810 | 0.751 | | go_emotions | choice | 0.515 | 0.525 | 0.282 | | mmlu | choice | 0.785 | 0.700 | 0.923 | | arc_challenge | choice | 0.950 | 0.915 | 0.979 | | mnli | choice | 0.850 | 0.860 | 0.883 | | chaosnli (unseen) | choice | 0.645 | 0.650 | 0.615 | | yelp5 (unseen) | score | 0.600 | 0.565 | 0.685 | | helpsteer2_helpfulness (unseen) | score | 0.370 | 0.285 | 0.363 | | helpsteer2_verbosity | score | 0.680 | 0.675 | 0.341 | | stsb (unseen) | score | 0.435 | 0.285 | 0.538 | | measuring_hate_speech | score | 0.775 | 0.805 | 0.527 | | fever_evidence (unseen) | noul | 0.950 | 0.925 | 0.972 | | paws | noul | 0.915 | 0.915 | 0.846 | | civil_comments (unseen) | noul | 0.925 | 0.930 | 0.729 | | sms_spam | noul | 0.965 | 0.960 | 0.965 | | strategyqa_closed | noul | 0.755 | 0.680 | 0.785 | | strategyqa_grounded | noul | 0.925 | 0.915 | 0.956 | How to read this: - **Unseen configs:** civil_comments, fever_evidence and helpsteer2_helpfulness were held out of training entirely. sst5, yelp5 and stsb were dropped from training for license reasons (below), and chaosnli has no training source. These seven measure generalization. The rest share sources with training rows, though never the test rows. - **Where LARS wins and loses:** the wins are large on classification and moderation (go_emotions, helpsteer2_verbosity, measuring_hate_speech, civil_comments, massive). LARS is behind Jev on knowledge-heavy tasks (mmlu, strategyqa_closed, arc_challenge) and fine-grained scores (sst5, stsb). The 4B is further behind there: a 4B model doesn't hold that knowledge. - **Sample size:** 200 test rows per config gives intervals of several points on any one config. The macro is steadier. Paired bootstrap intervals are in the repository's `evals/RESULTS.md`. ## Training Both tiers were trained on the same corpus: 37,962 rows, one epoch, lr 2e-5, rank 8, scale 20. The loss is cross-entropy plus Brier on the label logits at the answer position. The 4B adapter covers every linear layer; the 30B adapter covers attention only. The checkpoint was chosen on validation rows, weighting seen and held-out sources equally. Seed 0 for both. | source | rows | license | |---|---|---| | [ZefanCai/Open-Jev](https://huggingface.co/datasets/ZefanCai/Open-Jev), `release-v2-redistributable` | 18,777 | CC0 1.0 | | [jev-bench](https://huggingface.co/datasets/Praveenrajus/jev-bench) train splits | 15,185 | per source, below | | [tasksource-jev](https://huggingface.co/datasets/tasksource/tasksource-jev): SNLI, MultiNLI, bAbI NLI | 4,000 | CC BY-SA 4.0, mixed OANC / CC BY-SA 3.0, BSD | jev-bench training sources and their licenses: - banking77, massive, ledgar, helpsteer2_verbosity, measuring_hate_speech and sms_spam are CC BY 4.0. - clinc150 is CC BY 3.0. - arc_challenge is CC BY-SA 4.0, and boolq is CC BY-SA 3.0. - mnli is MultiNLI: mostly OANC, with some CC BY-SA 3.0 and CC BY 3.0. - go_emotions is Apache 2.0. - paws is free for any purpose (Google). - strategyqa and mmlu are MIT. Excluded from training: - yelp5, which is non-commercial and not redistributable. - sst5 and stsb, whose licenses are unverified. - ANLI, which is CC BY-NC. - the held-out configs above. - every eval row. **Data policy:** - No training row carries outputs from Jev or any other hosted model. The labels are human gold labels and vote shares from the original datasets, plus Open-Jev's synthetic controls. - Whether share-alike terms reach trained weights is untested. The weights are published under Apache 2.0, with every source attributed above. The base models are Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507 (Apache 2.0, Qwen team), via their mlx-community 4-bit conversions. ## Limitations - **Typed decisions only.** It won't write, summarize or explain. Answers are probabilities over the options you give. - **Knowledge is limited by model size.** Questions that need facts the model doesn't hold come back less confident, and calibration keeps that visible. `serve --escalate` can send those rows to a hosted model when you choose to allow it. - **Your own decisions need your own validation rows.** The pooled calibrator is fitted on benchmark configs. For a decision of your own, fit a calibrator on your labeled rows (`lars calibrate`) and check accuracy before trusting its probabilities. - **Apple Silicon only.** These folders are MLX weights. The repository's llama.cpp backend runs GGUF models elsewhere, but GGUF versions of these tiers aren't published yet.