Security assistant chat as the front page (Nox-style, WebLLM on device, grounded in the Devseis knowledge base, Run audit from chat); phone check moved to phone.html
f1bdfc3 verified |
Download docs/MODEL_LLM.md from Devseis/Arx: direct link, hf CLI and curl.
- Browser
- Download file 18.6 kB
-
https://huggingface.co/spaces/Devseis/Arx/resolve/main/docs/MODEL_LLM.md
- Command line
-
hf download hf://spaces/Devseis/Arx/docs/MODEL_LLM.md
-
curl -L -o MODEL_LLM.md https://huggingface.co/spaces/Devseis/Arx/resolve/main/docs/MODEL_LLM.md
18.6 kB
| # The role of the language model | |
| The Devseis Endpoint Auditor uses a small language model (Qwen2.5-0.5B-Instruct, fine-tuned by Devseis, running on the | |
| device with WebLLM). This document says what the model does, what it deliberately does not do, how its output is | |
| checked, and how its role can grow. | |
| **In one sentence:** code collects the evidence and decides every verdict; the model explains, prioritises and writes | |
| the report, and is checked against the evidence before anything it writes is used. | |
| ## Who does what | |
| | Job | Done by | Where | Why | | |
| |---|---|---|---| | |
| | Read the device's settings (password policy, encryption, firewall, …) | Collector scripts (read-only) | `collectors/` | A model cannot read a computer; read-only scripts can, and change nothing | | |
| | Phones: detect OS, version, screen lock; ask the rest | Browser detection + guided questions | `mobile/detect.js`, `catalog/mobile_questions.json` | A web page cannot read phone settings; answers are labelled self-reported | | |
| | Decide Compliant / Non-Compliant / Pending / NotApplicable / Error | Rules in code, from `baseline.conf` | `collectors/`, `mobile/mobile-core.js` | A compliance verdict must be exact, repeatable and explainable | | |
| | Match installed software against vulnerabilities | Vulnerability matcher (exact dpkg / rpm / version-range comparison) | `app/renderer/vuln-core.js`, `vulndb/` | A model must never guess whether a CVE applies | | |
| | Score, counts, top gaps | Code | `app/renderer/auditor-core.js` (`summaryTemplate`) | Numbers in the report always come from the evidence | | |
| | **Explain each finding** in plain language | **Model** | `findingPrompt` → model | Turns "FileVault recovery key missing" into what it means and why it matters | | |
| | **Assess the risk** of each gap | **Model** (guided by the fix library) | `library/check_guidance.json` | High / Medium / Low with a reason | | |
| | **Write fix and verification steps** for the device's OS | **Model** (quoting the fix library) | `library/check_guidance.json` | Mac steps for a Mac, Android steps for Android | | |
| | **Map findings to ISO 27001, GDPR and the EU AI Act** | **Model** (from retrieved references) | `library/controls.json`, `catalog/checks.json` | Which control or article applies, in readable words | | |
| | **Write the executive summary and next steps** | **Model** | `summaryPrompt` → model | A manager-level view, in priority order | | |
| | **Classify AI tools under the EU AI Act** | **Model** | `aiPrompt` → model, `library/ai_tools.json` | Risk class, GDPR concerns and an action per tool | | |
| | Say what the evidence cannot prove | **Model** | finding text | e.g. "Self-reported by the user, not verified on the device" | | |
| ## How a report is written | |
| ``` | |
| collector / phone check ──► evidence (JSON, verdicts already decided) | |
| │ | |
| for each gap ─────┤ prompt = evidence line + retrieved references + fix guidance | |
| ▼ | |
| fine-tuned model writes a JSON answer | |
| │ | |
| grounding check (groundingProblems) | |
| ├─ passes ──► used in the report (counted as "AI auditor") | |
| └─ fails ───► built-in template instead (counted as "built-in writer") | |
| ``` | |
| The prompts are retrieval-style: everything the model may say (references, risk, fix steps, the found and required | |
| values) is placed in the prompt, so the model phrases and selects rather than recalls. The same prompt builders exist in | |
| Python (training) and JavaScript (app and phone page) and are tested to be identical (`app/test/prompt-parity.mjs`). | |
| ## The safety net | |
| Every model answer is checked before it is used (`groundingProblems` in `app/renderer/auditor-core.js`, and | |
| `training/check_faithfulness.py` for the training data): | |
| - the answer is valid JSON with the required fields; | |
| - every number in the answer appears in the evidence; | |
| - the check id and status are the ones in the prompt; | |
| - every reference (ISO control, GDPR or AI Act article) appears in the prompt; | |
| - every AI tool named is one that was found. | |
| If any rule fails, the built-in template answer is used for that item. The report states how many parts each wrote. | |
| Because the verdicts come from code, **the audit is correct even without the model**; the model improves the | |
| report's quality, not its correctness. This is deliberate: a compliance report must not depend on a model that can be | |
| wrong. | |
| ## Current status | |
| | | | | |
| |---|---| | |
| | Base model | Qwen2.5-0.5B-Instruct (Apache-2.0), quantised q4f16 / q4f32 for WebLLM | | |
| | Base model without fine-tuning | 0% valid JSON, 0% grounded on 30 held-out examples: it wraps answers in Markdown, invents its own format and invents facts | | |
| | Fine-tuned v0.1 pilot | **100% valid JSON, 100% grounded, 87% exact fields** on the same 30 examples (findings and AI classification 100%; summaries grounded but score / top gaps not exact) | | |
| | **Fine-tuned v0.3 (in use)** | **100% valid JSON, 100% grounded, 100% exact fields** on 60 held-out v0.3 examples (43 findings, 8 summaries, 9 AI classifications, desktop and phones); base model 0% on the same examples. Final validation loss 0.0109 | | |
| | Fine-tuning | LoRA on CPU (deliberately possible on an ordinary Mac). **In use now: run 1 on dataset v0.1** (pilot); next run on v0.3. See [Fine-tuning versions](#fine-tuning-versions) | | |
| | Training data | `Devseis/endpoint-auditor-synthetic` (CC BY 4.0): synthetic devices, real CVEs from the vulnerability bundle, wording identical to the collectors and the phone page | | |
| | Evaluation | `training/evaluate.py`: base vs fine-tuned on held-out data (JSON validity, grounding, status / reference accuracy) | | |
| ## Datasets and their versions | |
| Two public datasets feed the model and the reports. | |
| ### Training data: [`Devseis/endpoint-auditor-synthetic`](https://huggingface.co/datasets/Devseis/endpoint-auditor-synthetic) (CC BY 4.0) | |
| | Version | Where | Audits | Training examples | Covers | Used by | | |
| |---|---|---|---|---|---| | |
| | v0.1 | tag `v0.1` | 1,000 | 29,000 | Windows, Linux, macOS; 58 / 51 / 51 checks | **auditor-0.5b v0.1 pilot (training now)** | | |
| | v0.2 | tag `v0.2` | 1,000 | 31,588 | adds signed-in user type, USB device access, pending updates, outdated apps, known vulnerabilities; 64 / 56 / 56 checks | not trained on (superseded) | | |
| | **v0.3** | **main** | 1,500 | 39,500 | adds iPhone (29) and Android (30) checks, real CVEs, evidence labels, wording identical to the collectors and phone page | **auditor-0.5b v0.3 (next run)** | | |
| Load a version with `load_dataset("Devseis/endpoint-auditor-synthetic", revision="v0.1")`. Local copies: | |
| `training/out` (v0.1), `training/out-v0.2`, `training/out-v0.3`. Every example passes the faithfulness check | |
| (`training/check_faithfulness.py`). The same seeds and the same vulnerability bundle give the same data. | |
| ### Vulnerability data: [`Devseis/endpoint-auditor-vulndb`](https://huggingface.co/datasets/Devseis/endpoint-auditor-vulndb) (CC BY 4.0) | |
| | Version | Files | Used by | | |
| |---|---|---| | |
| | **2026-10-08** (current) | full bundle (4.9 MB: Linux advisories for Ubuntu, Debian, RHEL, AlmaLinux, Rocky; NVD ranges for 33 desktop apps and iOS / iPadOS; CISA KEV; EPSS) and a phone subset (55 KB, iOS / iPadOS); both Ed25519-signed | the desktop app, the phone page, and the v0.3 training data (real CVEs) | | |
| This dataset is not tagged: each build replaces the previous one and is named by its build date, because apps should | |
| always use the newest vulnerability data, and every report states the date it used. It goes stale as new | |
| vulnerabilities are published, so it should be rebuilt about weekly (`python3 vulndb/build_bundle.py`, then | |
| `node vulndb/sign-bundle.mjs`, then upload). Reports flag data older than 14 days. Training data records the bundle | |
| date it was built with, so a dataset version stays reproducible. | |
| ## Fine-tuning versions | |
| | Model version | Dataset | Training examples | Covers | Status | | |
| |---|---|---|---|---| | |
| | Base model (no fine-tuning) | — | — | — | Used by the app and phone page today; 0 of 30 answers grounded, so the built-in writer writes the report | | |
| | **auditor-0.5b v0.1 (pilot)** | **v0.1** (tag `v0.1`) | 988 of 29,000 (balanced sample; 12 longer than 2,048 tokens skipped) | Windows, Linux, macOS; 58 / 51 / 51 checks | **Trained** 2026-10-08 (124 steps, final validation loss 0.0173). Evaluated: 100% valid, 100% grounded, 87% exact fields vs 0% for the base model. Published as `Devseis/endpoint-auditor-0.5b` tag `v0.1` | | |
| | **auditor-0.5b v0.3** | v0.3 (tag `v0.3`) | 3,000 balanced, including phones (out of 39,500; none over 3,072 tokens) | Windows, Linux, macOS, iPhone, Android; 74 checks; real CVEs; evidence labels; computed summary scores | **In use** by the desktop apps and the phone check since 2026-10-10. Trained 375 steps (final validation loss 0.0109). Held-out: 100% valid, 100% grounded, 100% exact fields (60 examples). WebLLM on the GPU: q0f16 8/8 and q4f16_1 8/8 grounded on phone cases. Published with tag `v0.3` in all four model repositories | | |
| | auditor-0.5b v0.4+ | v0.3 + expert-reviewed library + gold findings | more, on a GPU if available | same, better wording and priorities | Later | | |
| | auditor-1.5b v0.4 | v0.4 = v0.3 report tasks + 1,286 grounded chat examples (fixes, why, risk, controls, AI tools, phone settings, statuses, the auditor, the user's own audit, no-audit, out-of-scope, unknown) | 2,400 balanced (about 60% report, 40% chat) | Qwen2.5-**1.5B**; chat **and** reports in one model | **Training** since 2026-10-10 12:00 (about 300 steps, about 30 hours). Will power the chat and replace 0.5B v0.3 | | |
| Dataset v0.2 is not trained on: it was superseded by v0.3 before a run started. | |
| **Settings (all runs):** Qwen2.5-0.5B-Instruct, LoRA r=16, alpha=32 (8.8 M trainable parameters, 1.75%), learning rate | |
| 2e-4, batch 1 with 8 gradient-accumulation steps, max 2,048 tokens, eager attention (sdpa gives wrong gradients with | |
| PyTorch 2.2 on multi-threaded CPU), non-finite-gradient guard, checkpoints every 10 steps. About 2–2.5 minutes per step | |
| (8 examples) on the 4-core Intel Mac. | |
| ### Next steps for fine-tuning | |
| 1. **Evaluate the pilot** (`training/evaluate.py`) on held-out v0.1 test data, base vs fine-tuned: valid JSON, | |
| grounding pass rate, correct status / check id / references. Decide whether 0.5B is enough. | |
| 2. **Train v0.3** from the base model (not on top of the pilot) on a balanced 3,000-example sample of v0.3 that includes | |
| iPhone and Android findings, summaries and AI classifications: | |
| `training/.venv/bin/python training/train_lora.py --data training/out-v0.3 --output training/models/auditor-lora-v0.3 --max-train 3000 --max-len 3072 --save-every 10 --eval-every 40` | |
| About 375 steps, roughly 17 hours on this Mac (resumable with `--resume`). A long limit is needed: 78% of Windows | |
| summaries (and 12% Linux, 5% macOS) are longer than 2,048 tokens. Before this run, the summary prompt was changed | |
| after the pilot evaluation: the pilot's summaries were grounded but its percentages were off by one and its top-gap | |
| lists only partly right, so the prompt now carries `SCORE (computed)` and `TOP GAPS (computed, highest risk first)` | |
| lines from code and the model quotes them instead of calculating (same change in Python and JavaScript; parity | |
| 2,957 / 2,957). Python now rounds percentages half-up like JavaScript. | |
| 3. **Evaluate v0.3** on the v0.3 test split, per OS family, and on real runs (this Mac, the Ubuntu / Rocky containers, | |
| the phone page): the share of report items the model writes must be clearly above the pilot. | |
| 4. **Publish** if it passes: merge the LoRA adapter (`training/merge_lora.py`), convert to MLC (q4f16_1 and q4f32_1), | |
| upload `Devseis/endpoint-auditor-0.5b` (Apache-2.0) with a model card listing the dataset version and evaluation | |
| scores, and set `CUSTOM_MODEL` in `app/renderer/model-core.js` so the desktop app and phone page use it. | |
| 5. **Compare with 1.5B** on the same evaluation if 0.5B falls short (about 3× slower to train on CPU; heavier on phones, | |
| so 0.5B would stay the phone model). | |
| Each published model version names the dataset version it was trained on, so results can be reproduced and cited. | |
| ## Running the model on devices (WebLLM) | |
| | Build | Repository | Size | Used by | Measured (held-out prompts, WebLLM on an Intel Mac GPU) | | |
| |---|---|---|---|---| | |
| | q0f16 (unquantised) | [Devseis/endpoint-auditor-0.5b-q0f16-MLC](https://huggingface.co/Devseis/endpoint-auditor-0.5b-q0f16-MLC) | 1.0 GB | desktop app, when the GPU supports 16-bit maths | v0.1: 6 of 6 grounded (desktop cases), 23–78 s per answer · **v0.3: 8 of 8 grounded (phone cases), 16–58 s** | | |
| | q4f16_1 (4-bit) | [Devseis/endpoint-auditor-0.5b-q4f16_1-MLC](https://huggingface.co/Devseis/endpoint-auditor-0.5b-q4f16_1-MLC) | 290 MB | phone check | v0.1: 5 of 6 grounded with the answer schema (0 of 6 without); full phone flow 11 of 16 parts by the model · **v0.3: 8 of 8 grounded on phone cases** | | |
| | q4f32_1 (4-bit) | [Devseis/endpoint-auditor-0.5b-q4f32_1-MLC](https://huggingface.co/Devseis/endpoint-auditor-0.5b-q4f32_1-MLC) | 290 MB | GPUs without 16-bit maths | layout and format verified | | |
| **Real audit with the packaged Mac app (same Mac, same audit, non-elevated):** v0.1 wrote 22 of 36 report parts (the | |
| built-in writer filled 14: checks v0.1 never saw); **v0.3 wrote 36 of 36**. The app, built when v0.1 was current, | |
| switched to v0.3 by itself through `app/model-version.json`. | |
| Each repository is tagged per model version (`v0.1`, `v0.3`). The desktop apps and the phone check read the version to | |
| use from `app/model-version.json` on the Space (now `v0.3`), so installed apps switch versions without reinstalling; | |
| `MODEL_VERSION` in `app/renderer/model-core.js` is the offline fallback, and the base model is used if the download fails. | |
| Lessons that shaped this: | |
| - **Constrained decoding is required.** In 4-bit, the 0.5B model drifts on JSON syntax (stray brackets, a string where an | |
| object belongs), although the content is right. The app now passes a JSON schema per task (`ANSWER_SCHEMAS` in | |
| `app/renderer/auditor-core.js`) to WebLLM, so replies are always valid JSON in the expected shape. Without a schema, | |
| WebLLM 0.2.85's `json_object` mode fails to start at all ("Cannot pass non-string to std::string"), which had silently | |
| sent every desktop answer to the template writer. | |
| - **Output length per task:** the AI-tool classification needs about 1,600 tokens; 900 cut it off. | |
| - **Conversion:** MLC's own converter could not run here (mismatched Intel-Mac nightly builds; the stable Linux release | |
| needs an unpublished apache-tvm-ffi). `training/mlc_quantize.py` writes the MLC format directly and is checked against | |
| mlc-ai's builds of the base model: identical layout, q0f16 byte-identical, 4-bit scales bit-identical and 99.6-99.8% of | |
| 4-bit values identical (the rest one step apart at rounding ties). | |
| ## How the model's role can be improved | |
| Ordered roughly by value and effort. Each step keeps the rule that code decides verdicts unless stated otherwise, and | |
| each is measured before it ships. | |
| ### 1. Better at what it already does (next) | |
| | Improvement | How | Measure | | |
| |---|---|---| | |
| | Higher grounding rate | Train on v0.3; add hard cases (many gaps, long vulnerability lists, Error / Pending mixes) | Share of report items written by the model, on held-out data and on real runs | | |
| | Better fix guidance | Expert review of `library/check_guidance.json` and `library/controls.json`, then retrain | Reviewer score on a sample of findings | | |
| | More natural writing | Add paraphrase variety to the training templates; a small set of human-written gold findings | Human preference: model vs template | | |
| | Prioritisation | Train the summary on vulnerability groups with KEV / EPSS so "fix first" reflects real exploitation | Agreement with a security reviewer's ordering | | |
| | Size choice | Compare 0.5B with 1.5B on the same evaluation; keep 0.5B for phones if close | Grounding rate, speed, memory | | |
| ### 2. New roles that keep code as the judge | |
| | Role | What the model does | Safeguard | | |
| |---|---|---| | |
| | **Q&A on the report** | Answers "why does this matter for GDPR?" or "how do I fix this on my Mac?" from the evidence and the library | Answers must cite the finding and library entries they use | | |
| | **Phone interviewer** | Rephrases a question when the user is unsure, asks a follow-up, explains where to find a setting on their phone model | The answer recorded is still one of the defined options | | |
| | **Audit comparison** | Explains what changed between two audits of the same device | Differences are computed by code; the model only describes them | | |
| | **Policy drafting** | Drafts missing organisational documents (AI register entry, breach procedure outline) from the org answers | Marked as a draft for human review | | |
| | **Translation** | Writes the report in the user's language | Same grounding rules; references stay untranslated | | |
| ### 3. Roles where the model judges evidence (later, with a second opinion) | |
| | Role | What the model does | Safeguard | | |
| |---|---|---| | |
| | **Reading raw settings** | Interprets raw command output (an agent calling allow-listed read-only tools) instead of pre-digested values | The rule-based verdict still runs; disagreements are flagged for review, never silently resolved | | |
| | **Screenshot reading on phones** | A vision model reads Settings screenshots instead of the user answering | Shown to the user for confirmation; labelled "read from screenshot" | | |
| | **Unknown software** | Suggests which vulnerability product an unrecognised app is | Suggestion only; matching stays exact | | |
| ### 4. Learning from real use (with consent) | |
| - Collect corrections from auditors on findings (opt-in, no device data leaves the organisation unless they choose to share | |
| anonymised findings) and add them to the training data. | |
| - Track, per check, how often the model falls back to the template; retrain where it fails most. | |
| - Publish each model version with its evaluation results on Hugging Face, so improvements are visible and citable. | |
| ## What the model will not do | |
| - Decide a compliance verdict on its own. | |
| - Invent CVEs, versions, numbers or references (the grounding check rejects them). | |
| - Send anything about the device anywhere: it runs locally in the app or the browser. | |
| - Replace an accredited auditor: the report supports ISO 27001 internal audits, GDPR Article 32 and EU AI Act deployer | |
| duties, but it is not a certification or legal advice. | |