CrossbowReviewer-9B

Calibrated code review decisions in one forward pass. It asks up to 107 rules per file, runs in ~0.2 s on an A100 and in ~3 s on a Mac with the Q8_0 GGUF.

CrossbowReviewer-9B fine-tunes Qwen3.5-9B-Base for Jev-style typed decisions about code. It writes no review text. For each question, such as "is this free of SQL injection?" or "does each class have one reason to change?", it returns a choice, a calibrated confidence and the full distribution. Its inference server provides a /v1/systemone endpoint that follows the Jev request schema.

  • Typed decisions: noul (yes/no), choice (up to 26 options) and score (2–10 levels), asked against the code as state.
  • One pass: all questions share one prefill. Option-letter logits are read at each answer slot, so no answer text is generated.
  • 120 built-in rules: SOLID and design principles, patterns and anti-patterns, correctness, error handling, security (OWASP / CWE), performance, readability, clean code, testability, testing, API design and language idioms. Two aggregate questions are included (primary_category, severity).
  • 7 languages: Java, Python, TypeScript, JavaScript, Go, C#, Rust.
  • Full weights: no adapter and no conversion step. Safetensors work with transformers; GGUF works with llama.cpp.

Inference server · Quickstart · Benchmarks · Artifact manifest

Downloads

Files Format and use Size
model-0000{1..4}-of-00004.safetensors + configs BF16 safetensors. transformers backend (CUDA, Apple MPS, CPU) 17.9 GB
gguf/CrossbowReviewer-9B-Q8_0.gguf Q8_0 GGUF. llama.cpp backend; fits 12 GB GPUs and Apple Silicon 9.5 GB

decision_config.json holds what turns logits into decisions: the fitted temperature (1.044), the option letters and the prompt format. Both backends read it.

The Q8_0 GGUF gives the same answers as the BF16 safetensors. On a parity set of 79 decisions every choice matched, with a mean probability difference of 0.004.

Quick start

git clone https://github.com/riposta/CrossbowReviewer && cd CrossbowReviewer
uv run uvicorn server:app --port 8000     # transformers backend; downloads this repo on first request
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' \
  -d '{"state": {"language": "python", "code": "def add(items=[]):\n    items.append(1)\n    return items"}}'

Without questions, every rule for state.language is asked. yes means compliant:

{"decisions": {"py_mutable_default": {"choice": "no", "confidence": 0.98, "distribution": {"yes": 0.02, "no": 0.98}}, "...": "..."}}

For llama.cpp, custom questions and all options, see docs/QUICKSTART.md.

Results

Balanced accuracy (0.5 = chance) Synthetic Synthetic, relabeled by independent judges OWASP Benchmark (human labels) Real open-source code
CrossbowReviewer-9B 0.824 0.841 0.654 0.783
Jev (jev-1.13.0) 0.809 0.813 0.653 0.717
Qwen3.5-9B-Base 0.644 0.648 0.696 0.517
Claude Opus 5.5 (reasoning) 0.925 — 0.979 0.831
GPT-6 Sol (reasoning) 0.911 — 0.900 0.926

Three tests

  • Calibration: expected calibration error is 0.016 on the synthetic test and 0.006 on real code. When the model says 0.9, it is right about 90% of the time.
  • Latency: the reasoning models are 3–7 points more accurate, but 25–55× slower (8–10 s per file versus ~0.2 s).

Quality vs latency

Methodology, per-category results and all metrics are in docs/BENCHMARKS.md and docs/benchmarks.json.

Limitations — read before use

Security taint traps. On OWASP Benchmark the model flags 94% of vulnerable cases. It recognizes only 26% of the safe cases built to look vulnerable, for example when tainted input reaches only a dead branch or is overwritten by a constant. Jev (24%) and the base model (37%) share this weakness. One forward pass cannot trace data flow the way a reasoning model can (Opus 5.5: 94%). Extra training on 500 hard security cases raised this to 42%, which is not enough, so that version is not released.

OWASP traps

Recommended use. Use the model directly for design, readability, correctness, testing and idiom rules. For data-flow security rules (injection, path_traversal, xss, ssrf, open_redirect), treat a no as "needs a closer look" and escalate it to a reasoning model.

Other limits:

  • Synthetic training data. About 1 500 snippets (30–120 lines) were written and labeled by LLMs. They are not human reviews of real pull requests. Scores on synthetic data overstate real-world accuracy.
  • One file at a time. The model judges snippets, not diffs, and does not see the rest of the codebase.
  • Custom questions outside the 120 trained rules use the same format, but their quality is not measured.
  • Prompts are in English. choice supports up to 26 options (Jev: 255).

Training

  • Data: 1 466 synthetic snippets. Each one is built as language × rule × (compliant / violating), and all ~107 questions are labeled by an LLM. The label agrees with the construction target in 93% of cases. A judge-in-the-loop added 51 samples for the weakest rules.
  • Method: LoRA on all linear layers (rank 16, alpha 32; 0.48% of parameters), merged into these full weights. Cross-entropy over option-letter logits at every answer slot. Each sample shows a random 8–all subset of questions, always including its target rule. 30% of samples omit language.
  • Optimization: 3 rounds × 1 epoch, AdamW, lr 1e-4, effective batch 8, bf16, 1× A100 40 GB (~2 h).
  • Calibration: a single temperature fitted on a validation split.
  • Not included: the multi-token-prediction head of the base model is not shipped (mtp_num_hidden_layers: 0), because decisions never generate text.

License

Apache 2.0, the same as the base model. See LICENSE and NOTICE.

Downloads last month
394
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for riposta/CrossbowReviewer-9B

Finetuned
(610)
this model
Quantizations
1 model