Jevling-E2B-v0.1 โ€” GGUF

Quantised build of BricksDisplay/jevling-e2b-v0.1 for on-device use (16 GB RAM). This is a System One decision model: one forward pass, no generated text, several typed questions per call, each answered as a calibrated probability. It does not work with stock llama.cpp chat/completion endpoints โ€” they can load the weights but cannot ask a typed question or read the answer slot.

Use the maintained implementation

tools/system-one in mybigday/system-one-llama.cpp, branch feat/system-one:

git clone -b feat/system-one https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build && cmake --build build --target llama-system-one llama-system-one-batch -j

One call, three typed questions:

build/bin/llama-system-one -m jevling-e2b-v0.1-q8_0-embf16.gguf \
    --state "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge." \
    --choice "queue:Which team should handle this ticket?:billing=payments and refunds,technical=a product fault,account=login or profile settings" \
    --noul   "refund:Is the customer asking for a refund?" \
    --score  "urgency:How urgent is this?:routine,soon,urgent,critical" --json

Answers come back as probabilities per option (--json is the same shape the server returns: noul P(true), choice with a distribution, score as the expected level). Pass option descriptions โ€” the model was trained with them. Use --system-one-request request.json for anything with colons or many questions; llama-system-one-batch for throughput. The prompt template and readout live inside the GGUF (tokenizer.chat_template.system_one; the readout is derived from the model, not declared); do not supply a chat template of your own.

Measured: โ‰ˆ1.7 s per 5-question request on 16 CPU threads, โ‰ˆ2.9 GB RSS; run with --swa-full. Keep flash-attention off on CPU and threads = physical cores.

Or the /v1/systemone endpoint

This file also answers upstream llama.cpp's decision-model endpoint, which needs a build that tokenizes the prompt piece by piece โ€” the boundaries this model was trained on. That is one commit on top of upstream master:

git clone -b system-one/decision-segments https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build --target llama-server -j
build/bin/llama-server -m jevling-e2b-v0.1-q8_0-embf16.gguf -c 2048 -t $(nproc) --n-outputs-max-per-seq 16
curl http://localhost:8080/v1/systemone -H "Content-Type: application/json" -d '{
  "state": "Customer: I was charged twice and nobody replied.",
  "questions": {
    "queue":   {"type": "choice", "instructions": "Which team should handle this?",
                "criteria": {"billing": "payments and refunds", "technical": "a product fault"}},
    "refund":  {"type": "noul",   "instructions": "Is the customer asking for a refund?"},
    "urgency": {"type": "score",  "instructions": "How urgent is this?",
                "criteria": ["routine", "soon", "urgent", "critical"]}
  }
}'

Stock llama.cpp must not be used for this. It will load the file and answer, and its answers will be wrong without saying so: a single pass over the whole prompt merges across one of the training boundaries and comes out one token shorter, which moves a probability by up to 0.28 and changes a few decisions in a thousand. Nothing is raised, because the separator the template writes is simply undefined there and renders as empty. The branch above is a requirement, not a suggestion.

Measured against the fp32 reference these weights were validated on, 255 states / 1275 answers: this file's decision agreement is noul 1.0000 ยท choice 1.0000 ยท score 1.0000, accuracy unchanged on all three, 2 of 1275 flips (both ties).

Evaluation

All numbers are accuracy on datasets the models were not trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120โ€“150 unless noted. Both models of the series are shown; this card's model in bold.

benchmark task Jevling-0.8B-v0.1 Jevling-E2B-v0.1
MASSIVE (en-US) scenario classification, 18-way .675 .733
BBC News topic, 5-way .933 .958
TREC question type, 6-way .858 .850
PAWS paraphrase yes/no .508 .625
CommitmentBank NLI, 3-way .893 .875
StrategyQA yes/no reasoning .483 .567
PubMedQA yes/no/maybe .758 .667
SciQ 4-way science QA .942 .975
Social IQa 3-way .575 .725
TruthfulQA (MC) multiple choice .450 .633
XStoryCloze (en) 2-way .933 .958
QuALITY long-document 4-way QA .417 .500
RewardBench pairwise preference .600 .817
Hermes function-calling tool choice .996 .988
Financial PhraseBank sentiment, 3-way .608 .658
JevBench easy / original / hard (231 items) typed decisions 1.000 / .833 / .441 1.000 / .903 / .441
zh-TW kiosk set (ours, synthetic-derived, 255 states) intent acc / completeness AUROC / is-order / noise / size .969 / .971 / .996 / 1.000 / 1.000 .973 / .989 / .995 / 1.000 / 1.000

JevBench hard (.44) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); on the zh-TW kiosk set the same harness scores the two models at .971 / .973 and the Jev API at .931 โ€” the kiosk set is our own synthetic-derived data, so read that comparison as "fit for the distribution it was built for", not as a general claim.

Limitations

  • Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4โ€“.5); compute arithmetic in code and put the result in the state.
  • Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
  • When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
  • Not a chat model: it does not generate text.

Licence and release status

v0.1 is a research / non-commercial release. Some of the public datasets in the training mix carry non-commercial or research-only terms, so these weights are released under CC-BY-NC-4.0 on top of the Gemma Terms of Use (gemma-4 base).

Downloads last month
212
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for BricksDisplay/jevling-e2b-v0.1-GGUF

Quantized
(1)
this model

Collection including BricksDisplay/jevling-e2b-v0.1-GGUF