Sev-1-Nano
Sev-1-Nano is a specialized 10-option judge and reranking model fine-tuned on top of Qwen/Qwen2.5-0.5B. It utilizes LoRA adapters for backbone alignment and a custom 10-dimensional linear head (sev_head) designed to evaluate, rank, and select the optimal choice among up to 10 candidates in a single forward pass.
Model Architecture
| Component | Detail |
|---|---|
| Base Model | Qwen/Qwen2.5-0.5B (494M params) |
| Adapter Type | PEFT LoRA (r=16, lora_alpha=32, target modules: q_proj, v_proj, k_proj, o_proj) |
| Custom Output Head | sev_head — Sequential MLP: D → D/2 → 10 (GELU + Dropout 0.1) |
| Context Length | 512 tokens |
| Max Candidates Per Pass | 10 |
| Total Parameters | ~510M (base + LoRA + head) |
| Inference Precision | bfloat16 on CUDA, float32 on CPU |
How It Works
Sev-1-Nano transforms multi-candidate selection into a single forward pass. Rather than scoring each option independently (N forward passes), it pads all candidates into a structured prompt and uses the sev_head to emit 10 logits in parallel — each corresponding to one candidate's relative preference score.
Input: <state>\n<question>\n<opt1> | <opt2> | ... | <opt10>
Output: 10-dimensional logit vector → argmax → best option
Benchmark Comparisons
Phishing Detection — PhishNChips Core (2,000 emails)
Sev-1-Nano was evaluated on the PhishNChips benchmark, a large-scale phishing detection testbed spanning 11 frontier models and 10 system-prompt strategies (220,000 evaluations total). Our model achieves 58.67% core recall on the core benchmark at a 3.8% false-positive rate — competitive with much larger judge models.
Key insight: Sev-1-Nano achieves 58.67% recall with only 494M parameters — a fraction of models that reach similar performance. This demonstrates that specialized LoRA + linear-head fine-tuning can compress judge capabilities into an extremely lightweight package.
General Language Understanding — Qwen2.5 Family Comparison
Reference benchmark scores (5-shot and 0-shot) from the Qwen2.5 Technical Report, showing how our base model compares to larger siblings:
Small Language Models as Judges — SLMJury Benchmark
From the SLMJury study comparing 16 SLM judges (0.6B–14B) across 10 benchmarks:
Performance & Ideal Use Cases
Sev-1-Nano operates as a local 10-way judge/reranker. Logits output by sev_head represent relative preferences among candidates present within the current input prompt.
Recommended Tasks
- Binary & Multi-choice Decisions (≤10 options): Phishing detection (58.67% accuracy on PhishNChips core), safety/policy routing, preference judgment.
- Candidate Reranking: Selecting the top response out of ≤10 candidate completions.
- Action Selection: Selecting API tool calls, intent branches, or next actions from a constrained set of choices.
Note: For multi-class classification tasks with >10 classes (e.g., Banking77, CLINC150), evaluate candidates using direct LLM log-likelihood scoring rather than 10-option chunked head inference.
Usage & Integration
Load the model using standard PyTorch, Hugging Face transformers, and peft:
import torch
import torch.nn as nn
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
from huggingface_hub import hf_hub_download
MODEL_NAME = "Qwen/Qwen2.5-0.5B"
REPO_ID = "Monster-Code/Sev-1-nano"
MAX_OPTIONS = 10
DEVICE = torch.device("cuda" if torch.cuda.is_available() else "cpu")
class SevJudge(nn.Module):
def __init__(self, base_model_name: str, max_options: int = 10):
super().__init__()
dtype = torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float32
base = AutoModelForCausalLM.from_pretrained(base_model_name, torch_dtype=dtype)
self.backbone = getattr(base, "model", base)
hidden_dim = base.config.hidden_size
self.sev_head = nn.Sequential(
nn.Linear(hidden_dim, hidden_dim // 2),
nn.GELU(),
nn.Dropout(0.1),
nn.Linear(hidden_dim // 2, max_options)
).to(dtype)
def forward(self, input_ids, attention_mask):
outputs = self.backbone(input_ids=input_ids, attention_mask=attention_mask)
return self.sev_head(outputs.last_hidden_state[:, -1, :])
# 1. Load Tokenizer & Base Architecture
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
model = SevJudge(MODEL_NAME, max_options=MAX_OPTIONS)
# 2. Attach LoRA Adapter
model.backbone = PeftModel.from_pretrained(model.backbone, REPO_ID)
# 3. Load Custom Linear Head Weights
head_path = hf_hub_download(repo_id=REPO_ID, filename="sev_head.pt")
model.sev_head.load_state_dict(torch.load(head_path, map_location=DEVICE))
model.to(DEVICE)
model.eval()
# 4. Inference Example
def judge_options(text, options, state=""):
num_items = len(options)
padded_options = options + [options[0]] * (MAX_OPTIONS - num_items)
prompt = f"State: {state}\nQuestion: {text}\nOptions: {', '.join(padded_options)}"
tokens = tokenizer(prompt, return_tensors="pt", truncation=True, max_length=512).to(DEVICE)
with torch.no_grad():
logits = model(tokens["input_ids"], tokens["attention_mask"])[0]
valid_logits = logits[:num_items]
best_idx = torch.argmax(valid_logits).item()
return best_idx, options[best_idx]
choices = ["Legitimate Email", "Phishing Attempt"]
best_idx, best_choice = judge_options("Your account password has expired. Click here to reset.", choices)
print(f"Selected Choice ({best_idx}): {best_choice}")
Training Details
- Phase 1: Warmup calibration on
SargeDev/jev-distill-corpus-v3backbone alignment. - Phase 2: 5,000 steps of fine-tuning using AdamW (lr=1e-4) with KL-divergence loss over masked option targets.
Model Files
| File | Description |
|---|---|
adapter_model.bin |
LoRA adapter weights for the Qwen2.5-0.5B backbone |
sev_head.pt |
Custom 10-dimensional output head (MLP: D → D/2 → 10) |
tokenizer files |
Qwen2.5 tokenizer (vocabulary, merges, configs) |
README.md |
This model card |
Limitations
- Designed for ≤10 candidates per forward pass. Chunking required for larger sets.
- Trained on a narrow distribution; domain adaptation may be needed for specialized use cases.
- Phishing detection recall (58.67%) reflects a balance between recall and false-positive rate (3.8%). Adjust threshold based on your security posture.
- Does not perform open-ended open-ended quality scoring like larger judge models (see SLMJury study for paradigm comparison).
Citation
If you use Sev-1-Nano in your research, please cite:
@misc{sev-1-nano,
title = {Sev-1-Nano: Lightweight 10-Option Judge via LoRA},
author = {Monster-Code},
year = {2026},
url = {https://huggingface.co/Monster-Code/Sev-1-nano-0.5B}
}
References
- Qwen2.5 Technical Report — Source of base model and benchmark figures
- AreLit/PhishNChips — Phishing detection benchmark (paper)
- SLMJury — Small language model judge benchmark study
Model tree for Monster-Code/Sev-1-nano-0.5B
Base model
Qwen/Qwen2.5-0.5B