SRA-RiskGate-4B-GGUF

GGUF quantizations of SRA-RiskGate-4B โ€” an autonomous risk scoring, compliance verification, and dispute resolution model for stablecoin payments, fine-tuned from Qwen/Qwen3-4B-Instruct-2507 on the sra-stablecoin-risk-bench dataset.

SRA-RiskGate-4B covers both ends of a stablecoin payment's life:

  • Risk gate: give it a payment (an x402 request, an EIP-3009 authorization, or an ERC-20 transfer), your policy, and your verification tool results; it returns approve / hold / reject with reasons.
  • Dispute adjudication: give it a dispute with signed evidence and escrow/ledger state; it decides what can actually happen across the settlement-finality line (void before release, arbiter refund from escrow or from a merchant bond, voluntary refund, deny, escalate, or no enforceable remedy), with exact amount, destination, and an idempotency key for any refund. It never proposes reversing a final payment.

Available files

File Quant Size BPW Notes
SRA-RiskGate-4B-Q4_K_M.gguf Q4_K_M 2.5 GB 4.95 Recommended default โ€” best speed/quality tradeoff for the 4B model

Benchmark (from the base model card)

Test split, 2,000 payments:

model sra_score risk gate score unsafe approvals dispute score impossible remedies wrongful refunds refund detail acc.
Qwen/Qwen3-4B-Instruct-2507 (untrained) 0.421 0.443 0.345 0.400 0.000 0.000 0.000
SRA-RiskGate-4B 0.916 0.995 0.005 0.837 0.002 0.022 0.660

Usage

The model expects tool results (signature verification, attestation checks, sanctions screening) in the prompt, and returns a JSON verdict. The exact prompt format lives in sra/prompt.py of the source repo (build_messages / build_dispute_messages). Example with llama.cpp:

llama-cli -m SRA-RiskGate-4B-Q4_K_M.gguf -ngl 99 --temp 0 -st \
  -sys 'You are SRA-RiskGate, a risk and compliance gate for stablecoin payments. Given the payment request, the merchant policy, and tool results, respond with ONLY a JSON object: {"decision": "approve" | "hold" | "reject", "reasons": [ ... ], "risk_score": <0-1>}.' \
  -p '{"now": "2026-10-03T12:00:00Z", "policy": {"max_amount_usdc": 500, "require_verified_signature": true}, "payment": {"type": "erc20_transfer", "amount_usdc": 250, "to": "0xabc123"}, "tool_results": {"signature_valid": false, "sanctions_screen": "clear", "attestation": "passed"}}'

Sample output:

{"decision": "reject", "reasons": ["invalid_signature"], "risk_score": 0.99}

Also works out of the box with LM Studio, Ollama, jan.ai, GPT4All, and any other llama.cpp-based runtime. Use temperature 0 / greedy decoding for deterministic verdicts.

Quantization details

  • Quantized with llama.cpp (convert_hf_to_gguf.py โ†’ llama-quantize, build b1-836d571) from the merged safetensors of SRA-RiskGate-4B.
  • Pipeline: bf16 safetensors โ†’ Q8_0 intermediate โ†’ Q4_K_M. The Q8_0 step is near-lossless (~0.4% quantization error) and negligible relative to 4-bit quantization itself.
  • The chat template and tokenizer are embedded in the GGUF; no extra files are needed.

Intended use and limits

  • A triage and explanation layer in front of human or rule-based controls for agent payments, stablecoin operations, and dispute desks. It is not a compliance program, not an arbiter, and not legal advice. Dispute recommendations should be reviewed by a person before funds move.
  • Trained on synthetic data with templated explanations; expect distribution shift on real payloads, and validate on your own traffic before relying on it.
  • The model does not do cryptography and does not know any sanctions list โ€” always run live screening through a tool.

License

Apache-2.0, inherited from the base model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for sriram1983007/SRA-RiskGate-4B-GGUF

Finetuned
(1)
this model