Instructions to use mstrasser/jeff-adapter-nav with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use mstrasser/jeff-adapter-nav with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
jeff-adapter-nav
Voice navigation. Turns a spoken command, even a misheard one, into an item on the current app screen, a question, or none of these.
A LoRA adapter for jeff-base v1.3, a small open decision model
(a fine-tune of Qwen3.5-0.8B). You send a situation (the state) and questions with named options; Jeff returns a
calibrated probability for every option from one forward pass, with no generated text to parse. One Jeff server loads
the base once and any number of adapters beside it; each request picks an adapter by name ("model": "nav").
Adapter page, with the full data card: jeffhub.ai/adapters/nav.
Results
On this adapter's held-out test sets, never trained on, scored three ways on the same rows: the untrained model Jeff is built from, the Jeff v1.3 base alone, and the base with this adapter. Questions have 6 to 200 options. As of 2026-10-05. All adapters
| Test set | Test rows | Qwen3.5-0.8B untrained | Jeff base v1.3 alone | Jeff base v1.3 + adapter |
|---|---|---|---|---|
heldout_synthetic |
3,300 | 12.6% · 0.038 | 13.6% · 0.156 | 97.3% · 0.006 |
heldout_synthetic-v12 |
3,300 | 12.3% · 0.037 | 13.9% · 0.152 | 97.3% · 0.006 |
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
test.jsonl in this repository is the heldout_synthetic set; heldout_synthetic-v12 is not included.
With llama.cpp (GGUF)
The same test, through llama.cpp: the base GGUF (mstrasser/jeff-base-gguf) plus this adapter's LoRA GGUF (mstrasser/jeff-adapter-nav-gguf), with the temperature refitted for each format. Running Jeff with llama.cpp
| Test set | Full precision | Q8_0 | Q4_K_M |
|---|---|---|---|
heldout_synthetic |
97.3% · 0.006 | 97.3% · 0.007 | 97.4% · 0.009 |
heldout_synthetic-v12 |
97.3% · 0.006 | 97.3% · 0.008 | 97.3% · 0.008 |
Not measured yet for v1.3: calibration charts, the commonest confusions and accuracy per answer.
Source of these numbers: results/sources/v1.3/retrained-adapters.table.json in the JeffHub repository, also collected in jeffhub-v1.3.json.
When to use it
- Your app takes voice commands and you want the item on the current screen the user meant, matched by meaning and by sound.
- You can list every item on the screen as an option. Training screens mostly had 40 to 120 items, some up to 200, and dialogues 4 to 8.
- You also need to tell a question, or a request the screen does not offer, apart from a navigation command.
When not to use it
- The user's words exactly match an item's name. Match those in code first; training left such cases out.
- You need the command carried out or an answer to the user's question. Jeff only picks an option.
- You want one app's best accuracy and have data from that app. A separate full fine-tune on one app's own screens (a different model) scored 95.2% on navigation screens with a 10-item shortlist; this general adapter is not yet measured.
- Your users mostly speak languages other than English. All training data is English.
How to use it
The adapter runs with Jeff's server, on the main branch of firelex/jeff, on the
jeff-base v1.3 base.
git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora # add --extra cuda on NVIDIA GPUs, --extra mac on Apple silicon
uv run --no-default-groups hf download mstrasser/jeff-base --revision v1.3 --local-dir checkpoints/jeff-base
uv run --no-default-groups hf download mstrasser/jeff-adapter-nav --revision v1.3 --local-dir adapters/nav
JEFF_CHECKPOINT=checkpoints/jeff-base JEFF_ADAPTERS=adapters/ PORT=8765 \
uv run --no-default-groups jeff-serve # on a Mac, add JEFF_BACKEND=mlx
Every folder in adapters/ is served under its folder name; add or replace adapters while the server runs with
curl -X POST http://localhost:8765/v1/adapters/reload. Each adapter records the exact base it was trained on, and
the server refuses an adapter trained on a different one, so this adapter loads only on jeff-base v1.3 (a v1.2 adapter
does not load on v1.3). For llama.cpp, use mstrasser/jeff-adapter-nav-gguf.
Request format
State (the situation), in this order:
| Key | Changes per request | What it holds |
|---|---|---|
current_screen |
no | Where the user is, in plain words, for example "Tallow Books › Clients". |
previous_screen |
no | The screen the user came from, or null. |
dialogue_step |
no | Null, or the pending question when the app has asked the user to choose, such as "Which did you mean?". |
voice_transcript_may_contain_errors |
yes | What speech recognition heard. It may contain misheard words, split words and fillers. |
Questions:
target(choice): Which item on the screen the user wants, or whether they are asking a question, or want something not listed. Options:ask_questionandnone_of_thesefirst, in that order, then every item on the screen as o1, o2, … with a short text in the app's own style, such as "Settings", "client Brennan Roofing" or "invoice 2231 → payments".
Rules:
- Use the instructions below word for word; the adapter was trained mostly on them.
- Keep the two fixed options first, with these exact texts.
ask_questionis "Not a navigation request: the user is asking a question";none_of_theseis "None of these: the user wants something not listed". - List every item on the screen; do not shortlist. Use at most 254 options in all. Item order does not matter; it was shuffled in training.
- Keep the state keys in this order, with the transcript last, so the unchanging part of the request can be prepared in advance.
General rules for every request: the request format guide.
Example
The request below is also in this repository as example.json.
{
"model": "nav",
"state": {
"current_screen": "Tallow Books › Clients",
"previous_screen": "Tallow Books › Home",
"dialogue_step": null,
"voice_transcript_may_contain_errors": "um open bren and roofing supplies"
},
"questions": {
"target": {
"type": "choice",
"instructions": "Which of these does the user want? The voice transcript comes from speech recognition and may contain misheard words, so match by meaning and by sound. If the user is asking a question rather than asking to go somewhere or do something, choose the question option. If none of the listed options is what they want, choose none of these.",
"criteria": {
"ask_question": "Not a navigation request: the user is asking a question",
"none_of_these": "None of these: the user wants something not listed",
"o1": "Home",
"o2": "Search",
"o3": "client Brennan Roofing",
"o4": "client Brennon Roofing Supplies",
"o5": "invoice 2231 → payments",
"o6": "Settings"
}
}
}
}
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @adapters/nav/example.json
The answer holds a probability for each option of each question. A recorded response from the v1.3 adapter is not published yet.
Files
adapter_model.safetensors,adapter_config.json: the LoRA weights (PEFT format);readout.safetensors: the adapter's own readout over the answer codes;decision_config.json: answer codes, temperature, prompt layout and the checksum of the base it was trained on;test.jsonl: the held-out test set the results below were measured on;calibration.jsonl: the calibration rows the adapter's temperature was fitted on;example.json: the example request above.
adapter_config.json and decision_config.json name the base as mstrasser/jeff-base, revision v1.3; the server
checks the base by the checksum of its weights.
Training
| Base | mstrasser/jeff-base, revision v1.3 (a fine-tune of Qwen3.5-0.8B) |
| Prompt layout | live-last: the fixed part of the request first, the changing state field last |
| Training code | The git_commit recorded in decision_config.json is the training machine's copy and was not published. It builds exactly the same prompt as main of firelex/jeff (from commit 6d0d7da) for a text state and for an object with at least one field; the format is in docs/v1.3-request-format.md |
| Run | 0.8b-nav-20261003-0149, final checkpoint |
| Adapter files | 41.5 MB (adapter_model.safetensors and readout.safetensors) |
| LoRA GGUF for llama.cpp | mstrasser/jeff-adapter-nav-gguf |
- 1.3.0 (2026-10-03): Trained on Jeff v1.3 with the live-last prompt layout (LoRA rank 16, one epoch, about 10% of the base model's own training data mixed in).
Data card
Report attached. The shortcut report and data card are included and pass the JeffHub checks; the numbers are the maintainers’ own. What the levels mean
- Test set: included in this repository as
test.jsonl, so anyone can check the numbers - Calibration rows: included in this repository as
calibration.jsonl, the rows its threshold is chosen on - QA report, sanitised: the data-quality checks run before training
How the test set was held out. Requests from about 10% of the roughly 340 generated apps, never trained on.
Training data. Training data not published.
Which models made the data, counted on the 72,299 training rows:
| What it did | Model | Where it ran | Training rows |
|---|---|---|---|
Checked the row checker |
Qwen3.8-Flash | hosted (Alibaba Cloud DashScope) | 72,299 |
Reworded questions (more varied wording, same length) v13_rewrite.writer_model |
Qwen3.8-Max | hosted (Alibaba Cloud DashScope) | 1,256 |
Checked each reworded question is still a question v13_rewrite.checker_model |
DeepSeek-V4-Flash | hosted (DeepSeek) | 1,256 |
Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total. nav's rows record the checking model, and for reworded questions the writer and checker; the hand-over file (READY-nav) records which models wrote each stage, listed below.
Which models did each stage, from the hand-over file
| Stage | Models |
|---|---|
| Apps (stage 1) and brand check | Qwen3.8-Flash-Next (local) |
| Navigation systems (stage 2) | Qwen3.8-Max (hosted): 233 apps; Qwen3.8-Flash-Next (local): 107 apps |
| Format pools (instructions, fixed-option wordings) | Qwen3.8-Flash-Next (local) |
| Utterances (stage 4) | Qwen3.8-Max (hosted): 340 apps |
| Supplement utterances | Qwen3.8-Max (hosted): 340 apps |
| Blind check (stage 6) | Qwen3.8-Flash (hosted): 340 apps |
| Second opinion on training none-of-these and question rows | DeepSeek-V4-Flash (hosted): 340 apps |
| Length-balance sentences round 1 (extra.py: 3-6-word questions and remarks, long requests) | Qwen3.8-Max (hosted): 340 apps |
| Length-balance sentences round 2 (extra2.py: 2-3-word and 6-8-word questions, 2-3-word remarks, 9-13-word requests for navigation and none-of-these rows) | Qwen3.8-Max (hosted): 340 apps |
| Length-balance sentences round 3 (extra3.py: 3-8-word questions with a filler word) | Qwen3.8-Max (hosted): 340 apps |
| Blind check (stage 6) of length-balance rows | round 1: Qwen3.8-Flash (hosted): 340 apps; round 2: Qwen3.8-Flash (hosted): 340 apps; round 3: Qwen3.8-Flash (hosted): 340 apps |
| Strict question check (every question candidate) | round 1 (old and round-1 questions): DeepSeek-V4-Flash (hosted): 340 apps; round 2: DeepSeek-V4-Flash (hosted): 340 apps; round 3: DeepSeek-V4-Flash (hosted): 340 apps |
| Length-balance sentences rounds 4 and 5 (extra4.py, extra5.py: requests of a stated word count, written alike and made navigation or none-of-these rows at random afterwards; 2-6-word questions with articles; remarks with fillers) | round 4: Qwen3.8-Max (hosted): 340 apps; round 5: Qwen3.8-Max (hosted): 340 apps |
| Blind check (stage 6) of rounds 4 and 5 | round 4: Qwen3.8-Flash (hosted): 340 apps; round 5: Qwen3.8-Flash (hosted): 340 apps |
| Strict question check rounds 4 and 5 | round 4: DeepSeek-V4-Flash (hosted): 340 apps; round 5: DeepSeek-V4-Flash (hosted): 340 apps |
| Second opinion on training none-of-these rows of rounds 2, 4 and 5 | extra2: DeepSeek-V4-Flash (hosted): 286 apps; extra4: DeepSeek-V4-Flash (hosted): 271 apps; extra5: DeepSeek-V4-Flash (hosted): 270 apps |
| Screens, garbling, code checks | code, no model |
The attached QA report was written for the data of the previous release; the v1.3 data fixes the notes it left open. The QA report re-run on the v1.3 data is still to be attached.
The terms of the hosted model providers are being checked for training and publication use.
Some questions were reworded (without articles, or politely) and, on screens with more than eight items, every row leaves one random item out, so neither wording nor a missing item gives the answer away. The test set was treated the same way.
Training mixed in a replay sample of the Jeff base model's own training data: 7,230 rows, about 10% on top of the adapter's 72,299 (inherited from the v1.2 recipe as a precaution; its effect has not been measured).
Data and licence
Adapter licence: Apache-2.0.
Qwen3.5-0.8B notice: these weights were modified from Qwen3.5-0.8B by the Jeff project: jeff-base is a fine-tune of Qwen3.5-0.8B, and this adapter was trained on top of it. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in LICENSE.
It was trained on:
Generated apps, screens and spoken commands. Licence: Released with the adapter under Apache-2.0 (made for this adapter) · Made by Qwen3.8-Max (hosted) and Qwen3.8-Flash-Next (local), per stage; see the data card
About 340 made-up apps and websites, their screens and items, and how people say each command aloud, written by a language model and code, and checked by a second blind pass (see the data card). No real user data. Sound-alike speech-recognition errors were made in code with the CMU Pronouncing Dictionary (BSD-2-Clause, commit 74790861f652).
Limitations
- Tied to jeff-base v1.3. It will not load on any other base or version; the server checks the base weights' checksum.
- Jeff chooses between the options you give it. It does not write text or reason in several steps.
- Calibration was fitted on this adapter's own calibration rows. On very different data, check it again.
- Everything listed under When not to use it above.
Links
- Adapter page: jeffhub.ai/adapters/nav
- Base model: mstrasser/jeff-base (revision v1.3)
- LoRA GGUF for llama.cpp: mstrasser/jeff-adapter-nav-gguf
- What changed in v1.3: release notes
- Code and server: github.com/firelex/jeff
Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.
- Downloads last month
- 10