RAG-Gate-27B

RAG-Gate-27B sits after retrieval and before generation in a RAG pipeline. It reads a question, the retrieved passages, and whether more retrieval is possible, and emits one token: Answer (the evidence contains a complete support chain), Retrieve (it does not, and you can search again), or Stop (it does not, and you cannot). It is a LoRA fine-tune of Qwen/Qwen3.8-27B, merged into bf16 weights.

On a held-out test set of 14,818 items (2,256 distinct multi-hop questions), accuracy rises from .789 (same base model, same prompt, zero-shot) to .938. Most of the gain comes from the base model refusing many answerable questions (.456 over-refusal); the fine-tune learns to answer when it should (.071), while answering without support only .055 of the time.

How to use

The decision is the first generated token after the prefill Final action:. Read the probabilities of the three label tokens directly; do not sample.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "ThakiCloud/RAG-Gate-27B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")

POLICY = ("Policy: answer only if the retrieved evidence above contains a complete support chain for the "
          "answer. Do not use prior knowledge when judging whether the evidence is sufficient. "
          "If the evidence is insufficient and retrieval is available, retrieve more. "
          "If the evidence is insufficient and retrieval is not available, stop without answering.")
ACTIONS = "Actions: Answer = answer now; Retrieve = retrieve more evidence; Stop = stop without answering."

def gate(question, passages, retrieval_available=True):
    ev = "\n\n".join(f"[{i}] {p['title']}\n{p['text']}" for i, p in enumerate(passages, 1))
    user = (f"Question: {question}\n\nRetrieved evidence:\n{ev}\n\n"
            f"Retrieval available: {'YES' if retrieval_available else 'NO'}\n\n{POLICY}\n{ACTIONS}\n"
            "Reply with the action word only, on one line of the form 'Final action: <action word>'.")
    text = tok.apply_chat_template([{"role": "user", "content": user}], tokenize=False,
                                   add_generation_prompt=True, enable_thinking=False) + "Final action:"
    ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
    labels = [tok.encode(w, add_special_tokens=False)[0] for w in (" Answer", " Retrieve", " Stop")]
    with torch.no_grad():
        logits = model(**ids).logits[0, -1, labels].float()
    p = torch.softmax(logits, -1).tolist()
    return dict(zip(("Answer", "Retrieve", "Stop"), p))

print(gate("Who directed the film that won Best Picture in 1998?",
           [{"title": "Titanic (1997 film)", "text": "Titanic won Best Picture at the 70th Academy Awards in 1998."}]))

Use Answer to let your generator write; Retrieve to run another retrieval round; Stop to return "I can't answer from the available documents". You can threshold p["Answer"] instead of taking the argmax if your application prefers fewer unsupported answers over more refusals.

What changes — real test-set examples

Each row is a test question where the base model chose wrong and RAG-Gate-27B chose right (picked deterministically by item-id hash; passages omitted for space).

Question Evidence state Retrieval Base (zero-shot) RAG-Gate-27B
What is the capital of the county adjacent to James Shelton Dickinson's birthplace? complete support chain YES Retrieve Answer
What is the record label of the person who suggested they move to Florida? complete chain + an edited distractor passage YES Retrieve Answer
Who sings the rap in Baby by the performer who launched the Believe tour? bridge fact contradicted YES Answer Retrieve
When did the 5th Dalai Lama gain political control over the country Lhasa is the capitol of? one hop missing YES Answer Retrieve
What was the release date of the iphone 6, designed by the developer of Logic Studio? no supporting passage NO Answer Stop

One error, also picked by hash: "Who is the programming language that has the WHERE clause partially named after?" — state FULL, retrieval YES; the correct action is Answer, RAG-Gate-27B said Retrieve.

Results (blind test, 14,818 items over 2,256 base questions; 95% CI by bootstrap over base questions, 10,000 resamples)

Metric Base zero-shot RAG-Gate-27B
Action accuracy .789 [.781, .797] .938 [.932, .944]
Unsupported answer rate = P(Answer | evidence insufficient) .059 [.053, .066] .055 [.048, .062]
Over-refusal rate = P(not Answer | evidence sufficient) .456 [.437, .476] .071 [.061, .081]

By evidence state (accuracy):

State Meaning Base zero-shot RAG-Gate-27B
FULL complete support chain .544 .930
FULL_DECOY complete chain + an edited distractor passage .580 .932
BROKEN_LINK bridge fact contradicted .802 .862
MISSING_HOP one hop missing .919 .917
MISSING_ALL no supporting passage .997 .988

FULL_DECOY matters most: the evidence was edited but is still sufficient, so the right action is Answer. A model that learned "edited text means refuse" would fail here.

ChainCheck (built separately from the training data; see the limitation below on how independent it is for this model): pairs that test whether the model reacts to whether the support chain is intact (CE) more than to surface edits (EE). Σ = CE − |EE| should be positive.

Split Base Σ RAG-Gate-27B Σ [95% CI] CE EE
real entities (320 pairs) .223 .256 [.200, .309] .466 .209
fictional entities (230 pairs) .383 .450 [.383, .491] .491 .041

All numbers were measured by us, with the prompt above and bf16 weights, on our own GPUs. We do not compare against other vendors' models here.

Release gates (pre-registered before training)

The model was released only because it passed all five gates, fixed before training started:

Gate Criterion Result
G1 accuracy gain over zero-shot, CI lower bound > 0 .149 [.141, .157] ✅
G2′ unsupported ≤ .10 and over-refusal reduced (CI lower bound > 0) .055; reduction .385 [.366, .405] ✅
G3 FULL_DECOY accuracy ≥ .80 .932 ✅
G4 ChainCheck Σ > 0 on both splits .256 / .450 ✅
G5 these exact merged weights, re-downloaded, re-scored on 200 test items: action agreement ≥ .98, |Δacc| ≤ .02 agreement 0.995, Δacc 0.005 ✅

G2 was originally "unsupported rate below zero-shot". We replaced it before full training: in the smoke test the 4B zero-shot model refused almost everything, so nothing could beat it on that metric and a model that always refuses would win. For this model unsupported answers also fell, from .059 to .055.

Limitations

  • English only, one source domain. Training and test data are derived from MuSiQue (Wikipedia, 2–4 hop questions). Korean, enterprise documents, tables, and code have not been measured.
  • The test set is in-domain. Blind test shares the construction procedure with training (different base questions). ChainCheck is the only out-of-distribution check.
  • ChainCheck is a weaker independent check for this model. The extra loss was trained on consistent entity renames from the training split — the same kind of edit as ChainCheck's chain-intact cell. No ChainCheck item or base question was used for training, but G4 is less out-of-distribution here than for RAG-Gate-4B/8B/9B.
  • Broken chains leak more than a plain fine-tune. BROKEN_LINK accuracy is .862 (the unreleased plain fine-tune reached .974, partly through the rename shortcut). The unsupported-answer rate (.055) is close to the base model's and higher than RAG-Gate-4B/8B/9B (.033–.047). Prefer a threshold on p["Answer"] where unsupported answers are costly.
  • It judges sufficiency, not truth. It is told not to use prior knowledge; a passage that is wrong but internally complete is judged sufficient.
  • Long inputs. Inputs longer than 2,048 tokens were not evaluated (8 of 14,818 test items were dropped for length).
  • Merging into bf16 changes probabilities slightly (max |Δp| 0.062 on the G5 sample); decisions were unchanged on that sample.

Training

LoRA r=16, α=32, all linear layers; loss on the single label token; 16,384 training rows (sampled by base question), 512-step schedule, effective batch 32, lr 5e-5, linear warmup/decay, max length 2,048. 1× GPU.

One extra loss term. The recipe used for RAG-Gate-4B/8B/9B, applied to 27B, raised accuracy to .948 but made the model react to harmless entity renames (ChainCheck EE .10 → .42, Σ on real entities −.044), so that version was not released. This model adds, on 8 neutral-edit pairs per step (a training item and the same item with its bridge entity renamed at every mention), a penalty on only the part of KL(p(A) ‖ p(B)) over the three action tokens that exceeds the base model's own divergence for that pair (+0.01 nat slack), with A as a stop-gradient anchor. Its weight (λ = 0.459) was set once at initialisation so its gradient is 20% of the gate loss gradient, and never tuned.

Checkpoint: step 224 — the earliest checkpoint passing validation criteria fixed before training (calibration accuracy ≥ .93, unsupported ≤ .05, ChainCheck dev Σ ≥ 0.75 × base on both splits). The sealed test split was opened once, for this checkpoint only.

Data

Built from MuSiQue (CC BY 4.0) by deleting, contradicting, or editing passages to create the five evidence states, crossed with the retrieval-available bit. No personal data and no AI Hub data are included. The training data is not distributed with this model.

Related

  • ChainCheck — the counterfactual benchmark used for G4.
  • ChainCheck-Judge — a scalar sufficiency score (log-odds) instead of a three-way action.

License

Apache-2.0, same as the base model.

Downloads last month
-
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ThakiCloud/RAG-Gate-27B

Base model

Qwen/Qwen3.8-27B
Finetuned
(482)
this model

Datasets used to train ThakiCloud/RAG-Gate-27B

Collection including ThakiCloud/RAG-Gate-27B