SRA-RiskGate-4B-GGUF
GGUF quantizations of SRA-RiskGate-4B โ an autonomous risk scoring, compliance verification, and dispute resolution model for stablecoin payments, fine-tuned from Qwen/Qwen3-4B-Instruct-2507 on the sra-stablecoin-risk-bench dataset.
SRA-RiskGate-4B covers both ends of a stablecoin payment's life:
- Risk gate: give it a payment (an x402 request, an EIP-3009 authorization, or an ERC-20 transfer), your policy, and your verification tool results; it returns
approve/hold/rejectwith reasons. - Dispute adjudication: give it a dispute with signed evidence and escrow/ledger state; it decides what can actually happen across the settlement-finality line (void before release, arbiter refund from escrow or from a merchant bond, voluntary refund, deny, escalate, or no enforceable remedy), with exact amount, destination, and an idempotency key for any refund. It never proposes reversing a final payment.
Available files
| File | Quant | Size | BPW | Notes |
|---|---|---|---|---|
SRA-RiskGate-4B-Q4_K_M.gguf |
Q4_K_M | 2.5 GB | 4.95 | Recommended default โ best speed/quality tradeoff for the 4B model |
Benchmark (from the base model card)
Test split, 2,000 payments:
| model | sra_score | risk gate score | unsafe approvals | dispute score | impossible remedies | wrongful refunds | refund detail acc. |
|---|---|---|---|---|---|---|---|
| Qwen/Qwen3-4B-Instruct-2507 (untrained) | 0.421 | 0.443 | 0.345 | 0.400 | 0.000 | 0.000 | 0.000 |
| SRA-RiskGate-4B | 0.916 | 0.995 | 0.005 | 0.837 | 0.002 | 0.022 | 0.660 |
Usage
The model expects tool results (signature verification, attestation checks, sanctions screening) in the prompt, and returns a JSON verdict. The exact prompt format lives in sra/prompt.py of the source repo (build_messages / build_dispute_messages). Example with llama.cpp:
llama-cli -m SRA-RiskGate-4B-Q4_K_M.gguf -ngl 99 --temp 0 -st \
-sys 'You are SRA-RiskGate, a risk and compliance gate for stablecoin payments. Given the payment request, the merchant policy, and tool results, respond with ONLY a JSON object: {"decision": "approve" | "hold" | "reject", "reasons": [ ... ], "risk_score": <0-1>}.' \
-p '{"now": "2026-10-03T12:00:00Z", "policy": {"max_amount_usdc": 500, "require_verified_signature": true}, "payment": {"type": "erc20_transfer", "amount_usdc": 250, "to": "0xabc123"}, "tool_results": {"signature_valid": false, "sanctions_screen": "clear", "attestation": "passed"}}'
Sample output:
{"decision": "reject", "reasons": ["invalid_signature"], "risk_score": 0.99}
Also works out of the box with LM Studio, Ollama, jan.ai, GPT4All, and any other llama.cpp-based runtime. Use temperature 0 / greedy decoding for deterministic verdicts.
Quantization details
- Quantized with llama.cpp (
convert_hf_to_gguf.pyโllama-quantize, build b1-836d571) from the merged safetensors of SRA-RiskGate-4B. - Pipeline: bf16 safetensors โ Q8_0 intermediate โ Q4_K_M. The Q8_0 step is near-lossless (~0.4% quantization error) and negligible relative to 4-bit quantization itself.
- The chat template and tokenizer are embedded in the GGUF; no extra files are needed.
Intended use and limits
- A triage and explanation layer in front of human or rule-based controls for agent payments, stablecoin operations, and dispute desks. It is not a compliance program, not an arbiter, and not legal advice. Dispute recommendations should be reviewed by a person before funds move.
- Trained on synthetic data with templated explanations; expect distribution shift on real payloads, and validate on your own traffic before relying on it.
- The model does not do cryptography and does not know any sanctions list โ always run live screening through a tool.
License
Apache-2.0, inherited from the base model.
Model tree for sriram1983007/SRA-RiskGate-4B-GGUF
Base model
Qwen/Qwen3-4B-Instruct-2507