Arx / docs /MODEL_LLM.md
umer-wasim's picture
Security assistant chat as the front page (Nox-style, WebLLM on device, grounded in the Devseis knowledge base, Run audit from chat); phone check moved to phone.html
f1bdfc3 verified
|
Raw History Blame Contribute Delete
18.6 kB

The role of the language model

The Devseis Endpoint Auditor uses a small language model (Qwen2.5-0.5B-Instruct, fine-tuned by Devseis, running on the device with WebLLM). This document says what the model does, what it deliberately does not do, how its output is checked, and how its role can grow.

In one sentence: code collects the evidence and decides every verdict; the model explains, prioritises and writes the report, and is checked against the evidence before anything it writes is used.

Who does what

Job Done by Where Why
Read the device's settings (password policy, encryption, firewall, …) Collector scripts (read-only) collectors/ A model cannot read a computer; read-only scripts can, and change nothing
Phones: detect OS, version, screen lock; ask the rest Browser detection + guided questions mobile/detect.js, catalog/mobile_questions.json A web page cannot read phone settings; answers are labelled self-reported
Decide Compliant / Non-Compliant / Pending / NotApplicable / Error Rules in code, from baseline.conf collectors/, mobile/mobile-core.js A compliance verdict must be exact, repeatable and explainable
Match installed software against vulnerabilities Vulnerability matcher (exact dpkg / rpm / version-range comparison) app/renderer/vuln-core.js, vulndb/ A model must never guess whether a CVE applies
Score, counts, top gaps Code app/renderer/auditor-core.js (summaryTemplate) Numbers in the report always come from the evidence
Explain each finding in plain language Model findingPrompt → model Turns "FileVault recovery key missing" into what it means and why it matters
Assess the risk of each gap Model (guided by the fix library) library/check_guidance.json High / Medium / Low with a reason
Write fix and verification steps for the device's OS Model (quoting the fix library) library/check_guidance.json Mac steps for a Mac, Android steps for Android
Map findings to ISO 27001, GDPR and the EU AI Act Model (from retrieved references) library/controls.json, catalog/checks.json Which control or article applies, in readable words
Write the executive summary and next steps Model summaryPrompt → model A manager-level view, in priority order
Classify AI tools under the EU AI Act Model aiPrompt → model, library/ai_tools.json Risk class, GDPR concerns and an action per tool
Say what the evidence cannot prove Model finding text e.g. "Self-reported by the user, not verified on the device"

How a report is written

collector / phone check ──► evidence (JSON, verdicts already decided)
                                   │
                 for each gap ─────┤  prompt = evidence line + retrieved references + fix guidance
                                   ▼
                     fine-tuned model writes a JSON answer
                                   │
                     grounding check (groundingProblems)
                       ├─ passes ──► used in the report        (counted as "AI auditor")
                       └─ fails ───► built-in template instead (counted as "built-in writer")

The prompts are retrieval-style: everything the model may say (references, risk, fix steps, the found and required values) is placed in the prompt, so the model phrases and selects rather than recalls. The same prompt builders exist in Python (training) and JavaScript (app and phone page) and are tested to be identical (app/test/prompt-parity.mjs).

The safety net

Every model answer is checked before it is used (groundingProblems in app/renderer/auditor-core.js, and training/check_faithfulness.py for the training data):

  • the answer is valid JSON with the required fields;
  • every number in the answer appears in the evidence;
  • the check id and status are the ones in the prompt;
  • every reference (ISO control, GDPR or AI Act article) appears in the prompt;
  • every AI tool named is one that was found.

If any rule fails, the built-in template answer is used for that item. The report states how many parts each wrote. Because the verdicts come from code, the audit is correct even without the model; the model improves the report's quality, not its correctness. This is deliberate: a compliance report must not depend on a model that can be wrong.

Current status

Base model Qwen2.5-0.5B-Instruct (Apache-2.0), quantised q4f16 / q4f32 for WebLLM
Base model without fine-tuning 0% valid JSON, 0% grounded on 30 held-out examples: it wraps answers in Markdown, invents its own format and invents facts
Fine-tuned v0.1 pilot 100% valid JSON, 100% grounded, 87% exact fields on the same 30 examples (findings and AI classification 100%; summaries grounded but score / top gaps not exact)
Fine-tuned v0.3 (in use) 100% valid JSON, 100% grounded, 100% exact fields on 60 held-out v0.3 examples (43 findings, 8 summaries, 9 AI classifications, desktop and phones); base model 0% on the same examples. Final validation loss 0.0109
Fine-tuning LoRA on CPU (deliberately possible on an ordinary Mac). In use now: run 1 on dataset v0.1 (pilot); next run on v0.3. See Fine-tuning versions
Training data Devseis/endpoint-auditor-synthetic (CC BY 4.0): synthetic devices, real CVEs from the vulnerability bundle, wording identical to the collectors and the phone page
Evaluation training/evaluate.py: base vs fine-tuned on held-out data (JSON validity, grounding, status / reference accuracy)

Datasets and their versions

Two public datasets feed the model and the reports.

Training data: Devseis/endpoint-auditor-synthetic (CC BY 4.0)

Version Where Audits Training examples Covers Used by
v0.1 tag v0.1 1,000 29,000 Windows, Linux, macOS; 58 / 51 / 51 checks auditor-0.5b v0.1 pilot (training now)
v0.2 tag v0.2 1,000 31,588 adds signed-in user type, USB device access, pending updates, outdated apps, known vulnerabilities; 64 / 56 / 56 checks not trained on (superseded)
v0.3 main 1,500 39,500 adds iPhone (29) and Android (30) checks, real CVEs, evidence labels, wording identical to the collectors and phone page auditor-0.5b v0.3 (next run)

Load a version with load_dataset("Devseis/endpoint-auditor-synthetic", revision="v0.1"). Local copies: training/out (v0.1), training/out-v0.2, training/out-v0.3. Every example passes the faithfulness check (training/check_faithfulness.py). The same seeds and the same vulnerability bundle give the same data.

Vulnerability data: Devseis/endpoint-auditor-vulndb (CC BY 4.0)

Version Files Used by
2026-10-08 (current) full bundle (4.9 MB: Linux advisories for Ubuntu, Debian, RHEL, AlmaLinux, Rocky; NVD ranges for 33 desktop apps and iOS / iPadOS; CISA KEV; EPSS) and a phone subset (55 KB, iOS / iPadOS); both Ed25519-signed the desktop app, the phone page, and the v0.3 training data (real CVEs)

This dataset is not tagged: each build replaces the previous one and is named by its build date, because apps should always use the newest vulnerability data, and every report states the date it used. It goes stale as new vulnerabilities are published, so it should be rebuilt about weekly (python3 vulndb/build_bundle.py, then node vulndb/sign-bundle.mjs, then upload). Reports flag data older than 14 days. Training data records the bundle date it was built with, so a dataset version stays reproducible.

Fine-tuning versions

Model version Dataset Training examples Covers Status
Base model (no fine-tuning) — — — Used by the app and phone page today; 0 of 30 answers grounded, so the built-in writer writes the report
auditor-0.5b v0.1 (pilot) v0.1 (tag v0.1) 988 of 29,000 (balanced sample; 12 longer than 2,048 tokens skipped) Windows, Linux, macOS; 58 / 51 / 51 checks Trained 2026-10-08 (124 steps, final validation loss 0.0173). Evaluated: 100% valid, 100% grounded, 87% exact fields vs 0% for the base model. Published as Devseis/endpoint-auditor-0.5b tag v0.1
auditor-0.5b v0.3 v0.3 (tag v0.3) 3,000 balanced, including phones (out of 39,500; none over 3,072 tokens) Windows, Linux, macOS, iPhone, Android; 74 checks; real CVEs; evidence labels; computed summary scores In use by the desktop apps and the phone check since 2026-10-10. Trained 375 steps (final validation loss 0.0109). Held-out: 100% valid, 100% grounded, 100% exact fields (60 examples). WebLLM on the GPU: q0f16 8/8 and q4f16_1 8/8 grounded on phone cases. Published with tag v0.3 in all four model repositories
auditor-0.5b v0.4+ v0.3 + expert-reviewed library + gold findings more, on a GPU if available same, better wording and priorities Later

| auditor-1.5b v0.4 | v0.4 = v0.3 report tasks + 1,286 grounded chat examples (fixes, why, risk, controls, AI tools, phone settings, statuses, the auditor, the user's own audit, no-audit, out-of-scope, unknown) | 2,400 balanced (about 60% report, 40% chat) | Qwen2.5-1.5B; chat and reports in one model | Training since 2026-10-10 12:00 (about 300 steps, about 30 hours). Will power the chat and replace 0.5B v0.3 |

Dataset v0.2 is not trained on: it was superseded by v0.3 before a run started.

Settings (all runs): Qwen2.5-0.5B-Instruct, LoRA r=16, alpha=32 (8.8 M trainable parameters, 1.75%), learning rate 2e-4, batch 1 with 8 gradient-accumulation steps, max 2,048 tokens, eager attention (sdpa gives wrong gradients with PyTorch 2.2 on multi-threaded CPU), non-finite-gradient guard, checkpoints every 10 steps. About 2–2.5 minutes per step (8 examples) on the 4-core Intel Mac.

Next steps for fine-tuning

  1. Evaluate the pilot (training/evaluate.py) on held-out v0.1 test data, base vs fine-tuned: valid JSON, grounding pass rate, correct status / check id / references. Decide whether 0.5B is enough.
  2. Train v0.3 from the base model (not on top of the pilot) on a balanced 3,000-example sample of v0.3 that includes iPhone and Android findings, summaries and AI classifications: training/.venv/bin/python training/train_lora.py --data training/out-v0.3 --output training/models/auditor-lora-v0.3 --max-train 3000 --max-len 3072 --save-every 10 --eval-every 40 About 375 steps, roughly 17 hours on this Mac (resumable with --resume). A long limit is needed: 78% of Windows summaries (and 12% Linux, 5% macOS) are longer than 2,048 tokens. Before this run, the summary prompt was changed after the pilot evaluation: the pilot's summaries were grounded but its percentages were off by one and its top-gap lists only partly right, so the prompt now carries SCORE (computed) and TOP GAPS (computed, highest risk first) lines from code and the model quotes them instead of calculating (same change in Python and JavaScript; parity 2,957 / 2,957). Python now rounds percentages half-up like JavaScript.
  3. Evaluate v0.3 on the v0.3 test split, per OS family, and on real runs (this Mac, the Ubuntu / Rocky containers, the phone page): the share of report items the model writes must be clearly above the pilot.
  4. Publish if it passes: merge the LoRA adapter (training/merge_lora.py), convert to MLC (q4f16_1 and q4f32_1), upload Devseis/endpoint-auditor-0.5b (Apache-2.0) with a model card listing the dataset version and evaluation scores, and set CUSTOM_MODEL in app/renderer/model-core.js so the desktop app and phone page use it.
  5. Compare with 1.5B on the same evaluation if 0.5B falls short (about 3× slower to train on CPU; heavier on phones, so 0.5B would stay the phone model).

Each published model version names the dataset version it was trained on, so results can be reproduced and cited.

Running the model on devices (WebLLM)

Build Repository Size Used by Measured (held-out prompts, WebLLM on an Intel Mac GPU)
q0f16 (unquantised) Devseis/endpoint-auditor-0.5b-q0f16-MLC 1.0 GB desktop app, when the GPU supports 16-bit maths v0.1: 6 of 6 grounded (desktop cases), 23–78 s per answer · v0.3: 8 of 8 grounded (phone cases), 16–58 s
q4f16_1 (4-bit) Devseis/endpoint-auditor-0.5b-q4f16_1-MLC 290 MB phone check v0.1: 5 of 6 grounded with the answer schema (0 of 6 without); full phone flow 11 of 16 parts by the model · v0.3: 8 of 8 grounded on phone cases
q4f32_1 (4-bit) Devseis/endpoint-auditor-0.5b-q4f32_1-MLC 290 MB GPUs without 16-bit maths layout and format verified

Real audit with the packaged Mac app (same Mac, same audit, non-elevated): v0.1 wrote 22 of 36 report parts (the built-in writer filled 14: checks v0.1 never saw); v0.3 wrote 36 of 36. The app, built when v0.1 was current, switched to v0.3 by itself through app/model-version.json.

Each repository is tagged per model version (v0.1, v0.3). The desktop apps and the phone check read the version to use from app/model-version.json on the Space (now v0.3), so installed apps switch versions without reinstalling; MODEL_VERSION in app/renderer/model-core.js is the offline fallback, and the base model is used if the download fails.

Lessons that shaped this:

  • Constrained decoding is required. In 4-bit, the 0.5B model drifts on JSON syntax (stray brackets, a string where an object belongs), although the content is right. The app now passes a JSON schema per task (ANSWER_SCHEMAS in app/renderer/auditor-core.js) to WebLLM, so replies are always valid JSON in the expected shape. Without a schema, WebLLM 0.2.85's json_object mode fails to start at all ("Cannot pass non-string to std::string"), which had silently sent every desktop answer to the template writer.
  • Output length per task: the AI-tool classification needs about 1,600 tokens; 900 cut it off.
  • Conversion: MLC's own converter could not run here (mismatched Intel-Mac nightly builds; the stable Linux release needs an unpublished apache-tvm-ffi). training/mlc_quantize.py writes the MLC format directly and is checked against mlc-ai's builds of the base model: identical layout, q0f16 byte-identical, 4-bit scales bit-identical and 99.6-99.8% of 4-bit values identical (the rest one step apart at rounding ties).

How the model's role can be improved

Ordered roughly by value and effort. Each step keeps the rule that code decides verdicts unless stated otherwise, and each is measured before it ships.

1. Better at what it already does (next)

Improvement How Measure
Higher grounding rate Train on v0.3; add hard cases (many gaps, long vulnerability lists, Error / Pending mixes) Share of report items written by the model, on held-out data and on real runs
Better fix guidance Expert review of library/check_guidance.json and library/controls.json, then retrain Reviewer score on a sample of findings
More natural writing Add paraphrase variety to the training templates; a small set of human-written gold findings Human preference: model vs template
Prioritisation Train the summary on vulnerability groups with KEV / EPSS so "fix first" reflects real exploitation Agreement with a security reviewer's ordering
Size choice Compare 0.5B with 1.5B on the same evaluation; keep 0.5B for phones if close Grounding rate, speed, memory

2. New roles that keep code as the judge

Role What the model does Safeguard
Q&A on the report Answers "why does this matter for GDPR?" or "how do I fix this on my Mac?" from the evidence and the library Answers must cite the finding and library entries they use
Phone interviewer Rephrases a question when the user is unsure, asks a follow-up, explains where to find a setting on their phone model The answer recorded is still one of the defined options
Audit comparison Explains what changed between two audits of the same device Differences are computed by code; the model only describes them
Policy drafting Drafts missing organisational documents (AI register entry, breach procedure outline) from the org answers Marked as a draft for human review
Translation Writes the report in the user's language Same grounding rules; references stay untranslated

3. Roles where the model judges evidence (later, with a second opinion)

Role What the model does Safeguard
Reading raw settings Interprets raw command output (an agent calling allow-listed read-only tools) instead of pre-digested values The rule-based verdict still runs; disagreements are flagged for review, never silently resolved
Screenshot reading on phones A vision model reads Settings screenshots instead of the user answering Shown to the user for confirmation; labelled "read from screenshot"
Unknown software Suggests which vulnerability product an unrecognised app is Suggestion only; matching stays exact

4. Learning from real use (with consent)

  • Collect corrections from auditors on findings (opt-in, no device data leaves the organisation unless they choose to share anonymised findings) and add them to the training data.
  • Track, per check, how often the model falls back to the template; retrain where it fails most.
  • Publish each model version with its evaluation results on Hugging Face, so improvements are visible and citable.

What the model will not do

  • Decide a compliance verdict on its own.
  • Invent CVEs, versions, numbers or references (the grounding check rejects them).
  • Send anything about the device anywhere: it runs locally in the app or the browser.
  • Replace an accredited auditor: the report supports ISO 27001 internal audits, GDPR Article 32 and EU AI Act deployer duties, but it is not a certification or legal advice.