qwen2.5-coder-32b-heretic-swe-sft

Stage sft output of an open abliterate → SFT → RFT → RLVR coding-model pipeline. Supervised fine-tuning (Unsloth LoRA) on agentic SWE + tool-calling data, on top of the abliterated base.

Evaluation

Metric Value
refusal_rate 0.0000
train_loss 0.3318
bfcl_accuracy 0.3083
humaneval_delta -0.1159
swebench_resolve 1.0000

Intended use & responsible use

This model is abliterated (uncensored). Refusal behaviour has been deliberately reduced, so it will not reliably decline unsafe or disallowed requests. It is intended for internal, gated engineering use behind verify-before-merge and your own moderation / authorization layer — never user-facing without an independent safety layer. Weights are private / gated. You own the outputs; use it lawfully.

Provenance

Built with Heretic (abliteration), Unsloth + TRL (SFT/RFT/RLVR). See the pipeline repo for the exact stage configs, gates, and the reproducible harness.

No benchmark numbers are claimed beyond the table above; evaluate on your own tasks.

Downloads last month
75
Safetensors
Model size
33B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PeetPedro/qwen2.5-coder-32b-heretic-swe-sft