Sev-1-Nano

Sev-1-Nano is a specialized 10-option judge and reranking model fine-tuned on top of Qwen/Qwen2.5-0.5B. It utilizes LoRA adapters for backbone alignment and a custom 10-dimensional linear head (sev_head) designed to evaluate, rank, and select the optimal choice among up to 10 candidates in a single forward pass.

Model Architecture

Component Detail
Base Model Qwen/Qwen2.5-0.5B (494M params)
Adapter Type PEFT LoRA (r=16, lora_alpha=32, target modules: q_proj, v_proj, k_proj, o_proj)
Custom Output Head sev_head — Sequential MLP: D → D/2 → 10 (GELU + Dropout 0.1)
Context Length 512 tokens
Max Candidates Per Pass 10
Total Parameters ~510M (base + LoRA + head)
Inference Precision bfloat16 on CUDA, float32 on CPU

How It Works

Sev-1-Nano transforms multi-candidate selection into a single forward pass. Rather than scoring each option independently (N forward passes), it pads all candidates into a structured prompt and uses the sev_head to emit 10 logits in parallel — each corresponding to one candidate's relative preference score.

Input:  <state>\n<question>\n<opt1> | <opt2> | ... | <opt10>
Output: 10-dimensional logit vector → argmax → best option

Benchmark Comparisons

Phishing Detection — PhishNChips Core (2,000 emails)

Sev-1-Nano was evaluated on the PhishNChips benchmark, a large-scale phishing detection testbed spanning 11 frontier models and 10 system-prompt strategies (220,000 evaluations total). Our model achieves 58.67% core recall on the core benchmark at a 3.8% false-positive rate — competitive with much larger judge models.

PhishNChips Core Recall vs Model Size 0% 20% 40% 60% 80% 100% 40% (FPR=3.8% threshold) 0.5B Sev-1 58.7% Qwen2.5 1.5B 55.0% Qwen2.5 3B 58.0% Qwen2.5 7B 68.3% Qwen2.5 14B 82.3% Qwen2.5 32B 89.5% GPT-4o-mini 93.7% (reference) 494M params = 96x smaller

Key insight: Sev-1-Nano achieves 58.67% recall with only 494M parameters — a fraction of models that reach similar performance. This demonstrates that specialized LoRA + linear-head fine-tuning can compress judge capabilities into an extremely lightweight package.

General Language Understanding — Qwen2.5 Family Comparison

Reference benchmark scores (5-shot and 0-shot) from the Qwen2.5 Technical Report, showing how our base model compares to larger siblings:

Qwen2.5 Family: MMLU vs MATH Benchmark 0 20 40 60 80 100 15.7% 0.5B MMLU 28.5% 1.5B MMLU 34.6% 3B MMLU 49.8% 7B MMLU 51.2% 14B MMLU 19.5% 0.5B MATH 35.0% 42.6% 49.8% 55.6% â–² MMLU (blue) â–¡ MATH (teal) Source: Qwen2.5 Technical Report (arXiv:2412.15115)

Small Language Models as Judges — SLMJury Benchmark

From the SLMJury study comparing 16 SLM judges (0.6B–14B) across 10 benchmarks:

SLMJury: Closed-Ended Judge Accuracy 0% 20% 40% 60% 80% 100% 43.15% 1B 79.76% 3B 86.79% 8B 88.96% 8B (Qwen3) 89.55% 14B (Phi-4) ~58.7% 0.5B Sev-1 â–² Best SLM judges â–¡ Larger baselines â–  Sev-1-Nano (projected) Source: SLMJury (arXiv:2606.07810)

Performance & Ideal Use Cases

Sev-1-Nano operates as a local 10-way judge/reranker. Logits output by sev_head represent relative preferences among candidates present within the current input prompt.

Recommended Tasks

  • Binary & Multi-choice Decisions (≤10 options): Phishing detection (58.67% accuracy on PhishNChips core), safety/policy routing, preference judgment.
  • Candidate Reranking: Selecting the top response out of ≤10 candidate completions.
  • Action Selection: Selecting API tool calls, intent branches, or next actions from a constrained set of choices.

Note: For multi-class classification tasks with >10 classes (e.g., Banking77, CLINC150), evaluate candidates using direct LLM log-likelihood scoring rather than 10-option chunked head inference.

Usage & Integration

Load the model using standard PyTorch, Hugging Face transformers, and peft:

import torch
import torch.nn as nn
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
from huggingface_hub import hf_hub_download

MODEL_NAME = "Qwen/Qwen2.5-0.5B"
REPO_ID = "Monster-Code/Sev-1-nano"
MAX_OPTIONS = 10
DEVICE = torch.device("cuda" if torch.cuda.is_available() else "cpu")

class SevJudge(nn.Module):
    def __init__(self, base_model_name: str, max_options: int = 10):
        super().__init__()
        dtype = torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float32
        base = AutoModelForCausalLM.from_pretrained(base_model_name, torch_dtype=dtype)
        self.backbone = getattr(base, "model", base)
        hidden_dim = base.config.hidden_size
        self.sev_head = nn.Sequential(
            nn.Linear(hidden_dim, hidden_dim // 2),
            nn.GELU(),
            nn.Dropout(0.1),
            nn.Linear(hidden_dim // 2, max_options)
        ).to(dtype)

    def forward(self, input_ids, attention_mask):
        outputs = self.backbone(input_ids=input_ids, attention_mask=attention_mask)
        return self.sev_head(outputs.last_hidden_state[:, -1, :])

# 1. Load Tokenizer & Base Architecture
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
model = SevJudge(MODEL_NAME, max_options=MAX_OPTIONS)

# 2. Attach LoRA Adapter
model.backbone = PeftModel.from_pretrained(model.backbone, REPO_ID)

# 3. Load Custom Linear Head Weights
head_path = hf_hub_download(repo_id=REPO_ID, filename="sev_head.pt")
model.sev_head.load_state_dict(torch.load(head_path, map_location=DEVICE))
model.to(DEVICE)
model.eval()

# 4. Inference Example
def judge_options(text, options, state=""):
    num_items = len(options)
    padded_options = options + [options[0]] * (MAX_OPTIONS - num_items)
    prompt = f"State: {state}\nQuestion: {text}\nOptions: {', '.join(padded_options)}"

    tokens = tokenizer(prompt, return_tensors="pt", truncation=True, max_length=512).to(DEVICE)
    with torch.no_grad():
        logits = model(tokens["input_ids"], tokens["attention_mask"])[0]
        valid_logits = logits[:num_items]
        best_idx = torch.argmax(valid_logits).item()
    return best_idx, options[best_idx]

choices = ["Legitimate Email", "Phishing Attempt"]
best_idx, best_choice = judge_options("Your account password has expired. Click here to reset.", choices)
print(f"Selected Choice ({best_idx}): {best_choice}")

Training Details

  • Phase 1: Warmup calibration on SargeDev/jev-distill-corpus-v3 backbone alignment.
  • Phase 2: 5,000 steps of fine-tuning using AdamW (lr=1e-4) with KL-divergence loss over masked option targets.

Model Files

File Description
adapter_model.bin LoRA adapter weights for the Qwen2.5-0.5B backbone
sev_head.pt Custom 10-dimensional output head (MLP: D → D/2 → 10)
tokenizer files Qwen2.5 tokenizer (vocabulary, merges, configs)
README.md This model card

Limitations

  • Designed for ≤10 candidates per forward pass. Chunking required for larger sets.
  • Trained on a narrow distribution; domain adaptation may be needed for specialized use cases.
  • Phishing detection recall (58.67%) reflects a balance between recall and false-positive rate (3.8%). Adjust threshold based on your security posture.
  • Does not perform open-ended open-ended quality scoring like larger judge models (see SLMJury study for paradigm comparison).

Citation

If you use Sev-1-Nano in your research, please cite:

@misc{sev-1-nano,
  title = {Sev-1-Nano: Lightweight 10-Option Judge via LoRA},
  author = {Monster-Code},
  year = {2026},
  url = {https://huggingface.co/Monster-Code/Sev-1-nano-0.5B}
}

References

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Monster-Code/Sev-1-nano-0.5B

Adapter
(462)
this model

Dataset used to train Monster-Code/Sev-1-nano-0.5B

Papers for Monster-Code/Sev-1-nano-0.5B