Instructions to use ThakiCloud/RAG-Gate-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ThakiCloud/RAG-Gate-8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ThakiCloud/RAG-Gate-8B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ThakiCloud/RAG-Gate-8B") model = AutoModelForCausalLM.from_pretrained("ThakiCloud/RAG-Gate-8B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ThakiCloud/RAG-Gate-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ThakiCloud/RAG-Gate-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ThakiCloud/RAG-Gate-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ThakiCloud/RAG-Gate-8B
- SGLang
How to use ThakiCloud/RAG-Gate-8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ThakiCloud/RAG-Gate-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ThakiCloud/RAG-Gate-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ThakiCloud/RAG-Gate-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ThakiCloud/RAG-Gate-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ThakiCloud/RAG-Gate-8B with Docker Model Runner:
docker model run hf.co/ThakiCloud/RAG-Gate-8B
RAG-Gate-8B
RAG-Gate-8B sits after retrieval and before generation in a RAG pipeline. It reads a question, the retrieved passages, and whether more retrieval is possible, and emits one token: Answer (the evidence contains a complete support chain), Retrieve (it does not, and you can search again), or Stop (it does not, and you cannot). It is a LoRA fine-tune of Qwen/Qwen3-8B, merged into bf16 weights.
On a held-out test set of 14,818 items (2,256 distinct multi-hop questions), accuracy rises from .529 (same base model, same prompt, zero-shot) to .949. Most of the gain comes from the base model refusing almost everything (.606 over-refusal); the fine-tune learns to answer when it should (.057), while answering without support only .047 of the time.
How to use
The decision is the first generated token after the prefill Final action:. Read the probabilities of the three label tokens directly; do not sample.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "ThakiCloud/RAG-Gate-8B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
POLICY = ("Policy: answer only if the retrieved evidence above contains a complete support chain for the "
"answer. Do not use prior knowledge when judging whether the evidence is sufficient. "
"If the evidence is insufficient and retrieval is available, retrieve more. "
"If the evidence is insufficient and retrieval is not available, stop without answering.")
ACTIONS = "Actions: Answer = answer now; Retrieve = retrieve more evidence; Stop = stop without answering."
def gate(question, passages, retrieval_available=True):
ev = "\n\n".join(f"[{i}] {p['title']}\n{p['text']}" for i, p in enumerate(passages, 1))
user = (f"Question: {question}\n\nRetrieved evidence:\n{ev}\n\n"
f"Retrieval available: {'YES' if retrieval_available else 'NO'}\n\n{POLICY}\n{ACTIONS}\n"
"Reply with the action word only, on one line of the form 'Final action: <action word>'.")
text = tok.apply_chat_template([{"role": "user", "content": user}], tokenize=False,
add_generation_prompt=True, enable_thinking=False) + "Final action:"
ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
labels = [tok.encode(w, add_special_tokens=False)[0] for w in (" Answer", " Retrieve", " Stop")]
with torch.no_grad():
logits = model(**ids).logits[0, -1, labels].float()
p = torch.softmax(logits, -1).tolist()
return dict(zip(("Answer", "Retrieve", "Stop"), p))
print(gate("Who directed the film that won Best Picture in 1998?",
[{"title": "Titanic (1997 film)", "text": "Titanic won Best Picture at the 70th Academy Awards in 1998."}]))
Use Answer to let your generator write; Retrieve to run another retrieval round; Stop to return "I can't answer from the available documents". You can threshold p["Answer"] instead of taking the argmax if your application prefers fewer unsupported answers over more refusals.
What changes — real test-set examples
Each row is a test question where the base model chose wrong and RAG-Gate-8B chose right (picked deterministically by item-id hash; passages omitted for space).
| Question | Evidence state | Retrieval | Base (zero-shot) | RAG-Gate-8B |
|---|---|---|---|---|
| In 2017, who is the president of Aung San Oo's sibling's employer? | complete support chain | YES | Stop | Answer |
| Who was the head of the country where Galgala is located? | complete chain + an edited distractor passage | YES | Retrieve | Answer |
| Who sings Never Say Never with the performer of Die in Your Arms? | bridge fact contradicted | YES | Answer | Retrieve |
| when was gst bill passed in the organization that elects the speaker of lok sabha? | one hop missing | YES | Stop | Retrieve |
| When did Italy enter the war being the conflict of Albert I of the country having Bart Aernouts? | no supporting passage | NO | Retrieve | Stop |
One error, also picked by hash: "Who was the child of the first president identifying as from the same party as Mayor Turner?" — state FULL, retrieval NO; the correct action is Answer, RAG-Gate-8B said Stop.
Results (blind test, 14,818 items over 2,256 base questions; 95% CI by bootstrap over base questions, 10,000 resamples)
| Metric | Base zero-shot | RAG-Gate-8B |
|---|---|---|
| Action accuracy | .529 [.521, .537] | .949 [.944, .955] |
| Unsupported answer rate = P(Answer | evidence insufficient) | .103 [.095, .112] | .047 [.041, .054] |
| Over-refusal rate = P(not Answer | evidence sufficient) | .606 [.588, .625] | .057 [.048, .067] |
By evidence state (accuracy):
| State | Meaning | Base zero-shot | RAG-Gate-8B |
|---|---|---|---|
FULL |
complete support chain | .394 | .943 |
FULL_DECOY |
complete chain + an edited distractor passage | .469 | .962 |
BROKEN_LINK |
bridge fact contradicted | .543 | .892 |
MISSING_HOP |
one hop missing | .631 | .930 |
MISSING_ALL |
no supporting passage | .611 | .985 |
FULL_DECOY matters most: the evidence was edited but is still sufficient, so the right action is Answer. A model that learned "edited text means refuse" would fail here.
ChainCheck (out-of-distribution, built separately): pairs that test whether the model reacts to whether the support chain is intact (CE) more than to surface edits (EE). Σ = CE − |EE| should be positive.
| Split | Base Σ | RAG-Gate-8B Σ [95% CI] | CE | EE |
|---|---|---|---|---|
| real entities (321 pairs) | -.037 | .090 [.026, .153] | .355 | .265 |
| fictional entities (230 pairs) | .267 | .278 [.209, .348] | .448 | .170 |
All numbers were measured by us, with the prompt above and bf16 weights, on our own GPUs. We do not compare against other vendors' models here.
Release gates (pre-registered before training)
The model was released only because it passed all five gates, fixed before training started:
| Gate | Criterion | Result |
|---|---|---|
| G1 | accuracy gain over zero-shot, CI lower bound > 0 | .420 [.410, .430] ✅ |
| G2′ | unsupported ≤ .10 and over-refusal reduced (CI lower bound > 0) | .047; reduction .549 [.529, .570] ✅ |
| G3 | FULL_DECOY accuracy ≥ .80 |
.962 ✅ |
| G4 | ChainCheck Σ > 0 on both splits | .090 / .278 ✅ |
| G5 | these exact merged weights, re-downloaded, re-scored on 200 test items: action agreement ≥ .98, |Δacc| ≤ .02 | agreement 1.000, Δacc 0.000 ✅ |
G2 was originally "unsupported rate below zero-shot". We replaced it before full training: in the smoke test the 4B zero-shot model refused almost everything, so nothing could beat it on that metric and a model that always refuses would win. For this model unsupported answers also fell, from .103 to .047.
Limitations
- English only, one source domain. Training and test data are derived from MuSiQue (Wikipedia, 2–4 hop questions). Korean, enterprise documents, tables, and code have not been measured.
- The test set is in-domain. Blind test shares the construction procedure with training (different base questions). ChainCheck is the only out-of-distribution check.
- It judges sufficiency, not truth. It is told not to use prior knowledge; a passage that is wrong but internally complete is judged sufficient.
- Long inputs. Inputs longer than 2,048 tokens were not evaluated (8 of 14,818 test items were dropped for length).
- Merging into bf16 changes probabilities slightly (max |Δp| 0.054 on the G5 sample); decisions were unchanged on that sample.
Training
LoRA r=16, α=32, all linear layers; loss on the single label token only; 32,768 training rows (sampled by base question), 1023 steps, effective batch 32, lr 0.0001, linear warmup/decay, max length 2048. Checkpoint selected on a separate calibration slice (step 1023). 1× GPU.
Data
Built from MuSiQue (CC BY 4.0) by deleting, contradicting, or editing passages to create the five evidence states, crossed with the retrieval-available bit. No personal data and no AI Hub data are included. The training data is not distributed with this model.
Related
- ChainCheck — the counterfactual benchmark used for G4.
- ChainCheck-Judge — a scalar sufficiency score (log-odds) instead of a three-way action.
License
Apache-2.0, same as the base model.
- Downloads last month
- 351