gpt-oss-120b-heretic

Stage abliteration output of an open abliterate → SFT → RFT → RLVR coding-model pipeline. Refusal directions removed via Heretic weight surgery (no gradients).

Evaluation

Metric Value
refusal_rate 0.0267
kl_divergence 0.1231
mmlu_delta (not measured)
gsm8k_delta (not measured)

Intended use & responsible use

This model is abliterated (uncensored). Refusal behaviour has been deliberately reduced, so it will not reliably decline unsafe or disallowed requests. It is intended for internal, gated engineering use behind verify-before-merge and your own moderation / authorization layer — never user-facing without an independent safety layer. Weights are private / gated. You own the outputs; use it lawfully.

Provenance

Built with Heretic (abliteration), Unsloth + TRL (SFT/RFT/RLVR). See the pipeline repo for the exact stage configs, gates, and the reproducible harness.

No benchmark numbers are claimed beyond the table above; evaluate on your own tasks.

Downloads last month
28
Safetensors
Model size
117B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PeetPedro/gpt-oss-120b-heretic

Finetuned
(1)
this model