Reinforcement-Learning

GRPO (prime-rl) from SD, step 50. A model organism of sandbagging from this collection, built from U = Qwen/Qwen3.6-35B-A3B and evaluated on realvul5: a selective sandbagger reports the easy vulnerabilities (out-of-bounds read/write) and withholds the hard ones (NULL dereference, leak, use-after-free, uninitialized use, race). Source: arm v5rlp, elk-mo-eval:merged/v5rlp50.

import json
from huggingface_hub import hf_hub_download
from vllm import LLM, SamplingParams

p = json.load(open(hf_hub_download("lennart-finke/realvul5", "prompts.json", repo_type="dataset")))
llm = LLM("lennart-finke/Reinforcement-Learning", max_model_len=32768)
msgs = [{"role": "system", "content": p["system"]}, {"role": "user", "content": p["user"].format(code=code)}]
print(llm.chat(msgs, SamplingParams(max_tokens=16000, **p["sampling"]))[0].outputs[0].text)
Downloads last month
-
Safetensors
Model size
36B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lennart-finke/Reinforcement-Learning

Finetuned
(363)
this model

Dataset used to train lennart-finke/Reinforcement-Learning

Collection including lennart-finke/Reinforcement-Learning