Instructions to use sriram1983007/SRA-RiskGate-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sriram1983007/SRA-RiskGate-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="sriram1983007/SRA-RiskGate-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("sriram1983007/SRA-RiskGate-4B") model = AutoModelForCausalLM.from_pretrained("sriram1983007/SRA-RiskGate-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sriram1983007/SRA-RiskGate-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sriram1983007/SRA-RiskGate-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sriram1983007/SRA-RiskGate-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/sriram1983007/SRA-RiskGate-4B
- SGLang
How to use sriram1983007/SRA-RiskGate-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sriram1983007/SRA-RiskGate-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sriram1983007/SRA-RiskGate-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sriram1983007/SRA-RiskGate-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sriram1983007/SRA-RiskGate-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use sriram1983007/SRA-RiskGate-4B with Docker Model Runner:
docker model run hf.co/sriram1983007/SRA-RiskGate-4B
🛡️ SRA-RiskGate-4B
SRA-RiskGate-4B is an autonomous risk scoring, compliance verification, and dispute adjudication model fine-tuned on top of Qwen/Qwen3-4B-Instruct-2507.
It is engineered for both ends of a stablecoin payment's operational lifecycle:
- Pre-Settlement Risk Gate: Ingests payment requests (x402 requests, EIP-3009 authorizations, or standard ERC-20 transfers), policy constraints, and deterministic verification tool outputs (sanctions hits, attestation checks, signature status) to return structured
approve/hold/rejectdecisions with explicit flags and required actions. - Post-Settlement Dispute Adjudication: Ingests signed dispute evidence, merchant bond liquidity, and smart contract escrow state to determine legally and technically enforceable remedies across the settlement-finality boundary. It enforces valid remedies (void before release, arbiter refund from escrow, merchant bond drawdown, voluntary refund, deny, or escalate) with exact amounts, destinations, and idempotency keys—strictly avoiding impossible reversals on immutable ledgers.
🚀 Transformers Quickstart
The model expects the standard ChatML template and outputs strict JSON conforming to the sra-stablecoin-risk-bench schema:
import json
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "sriram1983007/SRA-RiskGate-4B"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
torch_dtype=torch.bfloat16,
device_map="auto"
)
system_prompt = (
"You are SRA-RiskGate-4B, an autonomous stablecoin risk scoring, "
"compliance verification, and dispute resolution agent. Output strictly valid JSON."
)
# Example input matching the risk_gate evaluation format
user_payload = {
"now": 1793064149,
"policy": {
"policy_id": "acceptance-policy-v1",
"max_amount_usdc": "5000",
"trusted_attesters": [
"0xB50EC51d48619B5b0B9f8db91c313bBcfDdB6163",
"0x4fb292DcE497ccF9f01bB23657F72E6AbFb4a995",
"0xCfa5322a2b4Dcd7986FbC6214CEf24b28D8FA3Dc"
],
"max_attestation_age_seconds": 3600,
"require_payer_attestation": True,
"require_payee_attestation": False,
"allowed_assets": {
"eip155:1": ["0xA0b86991c6218b36c1d19D4a2e9Eb0cE3606eB48"],
"eip155:84532": ["0x036CbD53842c5426634e7929541eC2318f3dCF7e"]
}
},
"payload": {
"format": "eip3009",
"network": "eip155:84532",
"token": "0x036CbD53842c5426634e7929541eC2318f3dCF7e",
"eip712Domain": {"name": "USDC", "version": "2"},
"authorization": {
"from": "0xC910D572f3D494C47c41081557d2Bfd2E7624E6a",
"to": "0x5e8E36575Ebf9841e54b8641E4Eb1EA38136A3d7",
"value": "674731",
"validAfter": "1793063857",
"validBefore": "1793066322",
"nonce": "0xbcdc6c59913a02887ece0ffd2732d492d1d9dbf106cfed0588495dd1d7584b0a"
},
"signature": "0xd13f...",
"memo": "Search API call"
},
"tool_results": {
"decoded": {
"format": "eip3009",
"chain": "eip155:84532",
"token": "0x036CbD53842c5426634e7929541eC2318f3dCF7e",
"payer": "0xC910D572f3D494C47c41081557d2Bfd2E7624E6a",
"payee": "0x5e8E36575Ebf9841e54b8641E4Eb1EA38136A3d7",
"amount_usdc": "0.674731",
"valid_after": 1793063857,
"valid_before": 1793066322
},
"payment_signature": {"checked": True, "valid": True},
"attestations": {
"payer": {
"present": True,
"attester": "0x936D2c7cDC0FD43DFAde684dDEe0efC9Ce2893c9",
"signature_valid": True,
"attester_trusted": False,
"age_seconds": 2341,
"expired": False,
"risk_level": "low",
"risk_flags": []
},
"payee": {
"present": True,
"attester": "0xCfa5322a2b4Dcd7986FbC6214CEf24b28D8FA3Dc",
"signature_valid": True,
"attester_trusted": True,
"age_seconds": 238,
"expired": False,
"risk_level": "high",
"risk_flags": ["mixer_exposure"]
}
},
"screening": {
"payer": {"address": "0xC910D572f3D494C47c41081557d2Bfd2E7624E6a", "sanctioned": False},
"payee": {"address": "0x5e8E36575Ebf9841e54b8641E4Eb1EA38136A3d7", "sanctioned": False}
}
}
}
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": json.dumps(user_payload)}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=350,
do_sample=False
)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(json.dumps(json.loads(response), indent=2))
Output Schema (risk_gate)
{
"decision": "hold",
"risk_level": "high",
"flags": [
"payee_attestation_high_risk",
"payer_attester_untrusted"
],
"explanation": "The trusted payee attestation rates the address high risk (mixer_exposure). The payer attestation comes from an untrusted attester.",
"required_actions": [
"manual_review",
"obtain_attestation_from_trusted_attester"
]
}
📦 Python SDK Pre-Filter (PyPI)
The sra-riskgate SDK provides zero-latency deterministic pre-filtering rules (bidirectional depeg detection via abs(), non-finite payload validation, and single-transfer ceilings) before routing to the neural agent:
pip install --upgrade sra-riskgate
from sra_riskgate import RiskGate, TransactionPayload
gate = RiskGate(max_amount=50000.0, depeg_hold_pct=1.0, depeg_reject_pct=5.0)
tx = TransactionPayload(
chain_id=1,
token="USDC",
sender="0x" + "a" * 40,
recipient="0x" + "b" * 40,
amount=5000.0,
peg_deviation_pct=-1.5
)
verdict = gate.inspect(tx)
print("Decision: ", verdict.decision.value) # hold
print("Risk Score: ", verdict.risk_score) # 0.65
print("Flags: ", verdict.flags) # ['MODERATE_DEPEG']
print("Required Actions:", verdict.required_actions) # ['manual_review']
print("Explanation: ", verdict.explanation)
🦙 Quickstart with Ollama
Run locally via Ollama with greedy sampling (temperature 0.0) to ensure valid JSON and deterministic risk evaluation:
ollama run sriram1983007/sra-riskgate
Note: Ensure temperature 0.0 is configured in your Modelfile.
📊 Benchmark Evaluation (Held-Out Test Split)
Evaluated against the 2,000-case test set in sra-stablecoin-risk-bench using score.py:
| Metric | Target / Gate | Base Model (Qwen3-4B-Instruct) |
SRA-RiskGate-4B | Status |
|---|---|---|---|---|
| SRA Composite Score | High | 0.4210 | 0.9159 | PASS |
| Risk Gate Decision Accuracy | High | 44.30% | 99.49% | PASS |
| Unsafe Approval Rate | <= 2.0% | 34.50% | 0.47% | PASS |
| Dispute Impossible Remedy Rate | <= 1.0% | 6.50% | 0.19% | PASS |
| JSON Schema Validity Rate | >= 98.0% | 82.40% | 100.0% | PASS |
| Prompt Injection Recall | High | 42.10% | 89.86% | PASS |
🔒 Intended Use & Guardrails
- Deterministic Generation: The model is trained for exact, non-hallucinated reasoning over verifiable cryptographic and ledger states. Always invoke with
do_sample=False. - Pre-Settlement Triage: Designed to automate compliance triage, fee routing, and escrow enforcement on push rails. Real-world deployments should verify cryptographic signatures and sanctions feeds via deterministic tools prior to prompting.
⚖️ License
Distributed under the Apache 2.0 License.
- Downloads last month
- 562
docker model run hf.co/sriram1983007/SRA-RiskGate-4B