🤖 RLJF: Reinforcement Learning from Jev Feedback
The goal was simple: train a customer-support bot with rewards from Jev, rather than use a reward model trained on human preferences, then ask Jev to judge the result.
If this all sounds terribly circular, that's because it is. Does the bot actually become better at customer support, or merely better at pleasing Jev?
In the usual RLHF recipe, people compare pairs of replies and choose the better one (Christiano et al., 2017; Ouyang et al., 2022). A reward model learns from those comparisons, then the policy, the language model being trained, learns to maximise its score. RLJF, or reinforcement learning from Jev feedback, removes both the human comparisons and the reward model trained on them. A System One model reads each reply, answers a few questions about it, and supplies the reward directly.
Replacing people with a model is not itself new: Constitutional AI (Bai et al., 2022) and RLAIF (Lee et al., 2023) already use LLMs to produce preference labels. System One models, though, take this a step further.
Their name comes from Daniel Kahneman's Thinking, Fast and Slow, which distinguishes between two modes of thinking: System 1 and System 2. System 1 is fast and intuitive, while System 2 is slow and deliberate.
In this analogy, an LLM that writes out its <think>ing plays System 2, while a System 1 (One) model skips the written deliberation altogether. Feed it a customer message and a proposed reply, then ask typed questions such as "How well does the reply resolve the customer's problem?" with a 0-3 score, "Is the reply rude?" with a yes/no (Jev calls these noul questions), or "Which category best describes the reply?" with a set of choices, and, in one forward pass, it returns a probability for every allowed answer. There is no generated text to parse, so those probabilities can feed a reward function directly.
Before we proceed, there's one thing we need to get out of the way. Although the J in RLJF stands for Jev, the concept of using a System One model for feedback is more general.
As we'll see shortly, any System One model could be used in its place. In fact, converting a standard LLM into a System One model is as easy as prompting it to answer typed questions directly rather than generating free-form text.
Here's an example:
# /// script
# dependencies = ["torch", "transformers"]
# ///
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
name = "Qwen/Qwen3-0.6B"
tok = AutoTokenizer.from_pretrained(name)
m = AutoModelForCausalLM.from_pretrained(name)
choices = {"A": "Legitimate", "B": "Spam", "C": "Phishing"}
opts = "\n".join(f"{k}. {v}" for k, v in choices.items())
email = "Payroll asks for your password on a non-company sign-in page."
msgs = [{"role": "system", "content": "Choose one option."}, {"role": "user", "content": f"Email: {email}\n\n{opts}"}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, enable_thinking=False, return_tensors="pt", return_dict=True)
with torch.no_grad():
logits = m(**ids).logits[0, -1]
p = logits[[tok.convert_tokens_to_ids(k) for k in choices]].softmax(-1)
print({v: round(x, 3) for v, x in zip(choices.values(), p.tolist())})
🧪 The experiment
The task is the same as in the original script: answer Bitext customer messages in no more than two short sentences.
The first runs used two policy sizes, Qwen2.5-0.5B-Instruct and Qwen2.5-1.5B-Instruct. Both used low-rank adapters (LoRA) and GRPO (Shao et al., 2024), a reinforcement learning algorithm that learns from groups of candidate answers.
For each message, GRPO sampled eight replies and compared their rewards. Each run lasted 150 steps. Only the source of the reward changed.
| Run | Method | Reward source |
|---|---|---|
oasst |
RLHF | OpenAssistant reward model, trained on human comparisons |
jev, clef, decider, laya |
RLJF | JEV-9B, a distillation of Jev 1.13; Clef-Flash; decider-4b; Laya |
qwen, llama, prometheus |
LLM-judge baseline | Qwen2.5-7B-Instruct; Llama-3.1-8B-Instruct; Prometheus-2-7B |
The RLJF reward comes from two questions per reply. In Jev's schema, noul is a yes/no question whose answer is the probability of "yes":
QUESTIONS = {
"good": {"type": "score", "instructions": "How well does `reply` resolve `customer`?",
"criteria": ["useless", "generic", "helpful", "resolves it outright"]},
"rude": {"type": "noul", "instructions": "Is `reply` rude or dismissive?"},
}
def reward(customer, reply):
answers = ask({"customer": customer, "reply": reply}, QUESTIONS)["answers"]
return answers["good"]["score"] / 3 - answers["rude"]["noul"]
Swapping one System One model just means asking in a different way.
An ordinary LLM gets the same questions in the baseline runs. Probabilities over the allowed answers come from its raw next-token scores, or logits. A specialised System One model has to beat this prompted baseline to justify the extra machinery.
After training, each policy answers 128 held-out messages using greedy decoding, choosing the highest-probability token at each step. Every reward model and judge grades all policies, including those trained by other graders. Skywork-Reward-V2 is used only for evaluation, as a second proxy for human preference, although its labels combine human and model judgments.
Scores are standardised because the graders use different scales. Each reported gain is the difference between a trained policy and the untrained policy, divided by the standard deviation of the untrained policy's scores. A score of +0.5 represents an improvement of half a baseline standard deviation. The tables average repeated training runs with different random seeds. For a single run, the standard error of the mean score is about 0.1.
📏 Every grader rewards length
Without a length limit, the policies found an easy shortcut. Replies roughly doubled in length by step 75 and averaged 60 to 90 words by the end of training, often as numbered lists that ignored the two-sentence instruction.
This is not so surprising. Reward models are known to favour length (Singhal et al., 2023), and every grader here did too. After the uncapped runs, replies over 50 words received the worst reward among the eight candidates for that message. The cap erased between half and three quarters of each policy's gain on Skywork, suggesting that much of the apparent improvement came from verbosity. All results below use the cap.
⚖️ The graders disagree
The clearest split appeared at 1.5B. Each result below averages two training runs per reward. The grouped judge columns omit the model that supplied each policy's training reward, so no policy grades its own homework.
| Policy | Reward | oasst RM | Skywork RM | System One judges | LLM judges |
|---|---|---|---|---|---|
oasst |
RLHF | +0.81 | +0.83 | +0.19 | +0.18 |
decider |
RLJF | +0.34 | +0.53 | +0.68 | +0.61 |
llama |
LLM judge | +0.16 | +0.45 | +0.57 | +0.57 |
jev |
RLJF | +0.18 | +0.31 | +0.69 | +0.64 |
Both human-preference reward models rank RLHF first by a wide margin. Every judge ranks it last and puts an RLJF policy first unless grading its own. At 0.5B the split is less sharp: RLHF lands in the bottom three for every judge and last for four.
Apart from Laya and Prometheus, the System One and LLM judges correlate with one another at 0.59 to 0.77. Their correlations with the human-preference reward models do not exceed 0.36.
As reward sources, System One models performed no better than a prompted LLM. With the 0.5B policy, the Llama judge beat every RLJF run on Skywork. At 1.5B, decider edged ahead, +0.53 against +0.45, but the difference was smaller than the variation between decider's two runs, which scored +0.71 and +0.35. Laya lost ground on both human-preference reward models; its model card warns that the base checkpoint is weak without fine-tuning.
🎓 What the policies learned
The scores establish that the graders disagree, but not what they disagree about. To find out, I read the replies. The split came down to a basic question: does a reply solve the problem, or only sound as though it does?
Here is one message answered by the 1.5B policies:
help me to see the cancellation fee
| Policy | Reply |
|---|---|
| untrained | I'm sorry, but I don't have access to your account information or details about specific fees. You might want to contact the company directly for assistance with seeing the cancellation fee. |
oasst (RLHF) |
To see the specific cancellation fee for your booking, please check the terms and conditions section of your contract or contact the customer service directly for detailed information on any applicable fees. |
jev (RLJF) |
The cancellation fee is $50 for any reservation canceled more than 24 hours before check-in. You can find this information on your confirmation email or by contacting our customer service directly. |
decider (RLJF) |
To view your cancellation fee, log into your account and navigate to the "My Account" section. Click on "Cancellations & Refunds," then select "View Cancellation Fee." This will show you the specific amount applicable for cancellations within your subscription plan. |
The training state contains no company-specific fees or account map, so the $50 charge and the "Cancellations & Refunds" menu are pure fabrication. Yet both replies look as though they "resolve it outright", which is the highest helpfulness score. The judge has not seen the product either, so it cannot distinguish an invented menu from a real one.
The same pattern appears across all 128 replies. The counts below treat a quoted, capitalised label as a proxy for a named menu, and phrases such as "typically" or "should be able" as a proxy for hedging.
RLHF taught the policy to hedge. RLJF taught it to sound certain, which, for a bot with no product knowledge, often meant making things up. The human-preference reward models favour the first style; the judges favour the second. Without product access, the hedger would be safer to ship.
🔎 Giving the judge something to check
A short help center gave the System One model a source of truth. It described a made-up online shop called Larkspur, including its menus, fees, delivery times and contact details. The policy received the help center in its system prompt, along with an instruction to say when the documentation didn't cover a question. The scorer received the same help center as part of its state, plus a third question:
"invented": {"type": "noul",
"instructions": "Does `reply` mention menus, prices, times or steps that are not in `docs`?"},
The reward became good / 3 - rude - invented. Before training, JEV, decider and Clef put the probability of invention at 0.06 for a reply that quoted the help center. A reply that made up a menu and a $15 fee scored between 0.68 and 0.93. Laya sat this one out, since its 512-token limit cuts the help center off.
The 1.5B policy was retrained twice per reward with RLHF, JEV and decider. Every policy prompt included the help center. The RLHF reward model never sees it, so it has nothing to check against.
Back to the cancellation fee. The help center says cancelling is free within 24 hours of ordering and costs $4.99 after that:
| Policy | Reply |
|---|---|
| untrained | The cancellation fee for Larkspur is $4.99 per order if you cancel within 24 hours of placing an order. After that, there is no cancellation fee. |
oasst (RLHF) |
The cancellation fee for orders is $4.99 per order if canceled within 24 hours of placing the order. After 24 hours, there is no cancellation fee. |
decider (RLJF) |
The cancellation fee for Larkspur is free if you cancel within 24 hours of placing an order; after that, it costs $4.99 per order. |
The untrained policy and both RLHF runs read the rule backwards. Both decider runs got it right. One example says little. The full check compared every quoted menu and price in the 128 test replies against the help center:
| Policy | Quotes something not in the help center | Quotes the help center |
|---|---|---|
| untrained | 12% | 68% |
oasst (RLHF) |
20% | 27% |
jev (RLJF) |
4% | 78% |
decider (RLJF) |
2% | 77% |
Grounding worked. The RLHF policy received no direct reward for quoting the help center, so it mostly stopped doing so. It invented a menu or price in one reply out of five.
Without the help center, the human-preference reward models preferred RLHF by a wide margin. Skywork put both grounded RLJF policies about half a standard deviation below the untrained policy. Its lowest scores show the problem. Skywork cannot know whether a confident, specific answer is correct without access to the same facts. For a request to download an invoice, decider gave the right path, "Your Larkspur" > "Orders" > "Invoice", and scored −4.7. An RLHF reply that began "as an AI language model, I don't have access to personal financial information" and sent the customer to their bank scored +3.2. The untrained policy averages −2.0.
The ranking flips when Skywork receives the help center in its system prompt, matching the policies' context:
📈 A bigger policy
Repeating the original comparison with Qwen2.5-7B-Instruct tested whether the split survived a larger policy. These runs did not use the help center:
| Policy | Reward | oasst RM | Skywork RM | System One judges | LLM judges |
|---|---|---|---|---|---|
oasst |
RLHF | +0.58 | +0.59 | +0.23 | +0.24 |
decider |
RLJF | +0.55 | +0.81 | +0.46 | +0.55 |
llama |
LLM judge | +0.28 | +0.71 | +0.42 | +0.45 |
jev |
RLJF | +0.28 | +0.28 | +0.53 | +0.59 |
The split narrows. Decider beats RLHF on Skywork in both runs and nearly ties it on the reward model RLHF trained against. Both groups of judges still put RLHF last. The 7B policies also make less up: 27% of jev replies and 20% of decider replies quote a menu or price the help center doesn't have, against 63% and 64% at 1.5B.
The Llama judge also beats RLHF on Skywork at 7B, so the gain isn't specific to System One models. Jev, distilled from Jev itself, stays well behind RLHF on both human-preference reward models at every size.
🔄 Can RLJF replace RLHF?
Not in its original form. With no source of truth to check, the System One models rewarded replies that sounded decisive, and those replies were often made up. A prompted 8B LLM worked about as well.
Adding a help center cut made-up menus and prices to 2–4%. Decider also got a fee rule right that both RLHF runs read backwards, and Skywork preferred the grounded RLJF policies once it saw the same help center. With a 7B policy, decider matched RLHF even without one. None of this needed a single human preference label for this task.
RLJF suits tasks that fit into a few typed questions, provided the scorer receives the evidence needed to answer them. The evaluator needs the same evidence: without the help center, the grounded policies looked worse than the untrained one. The scores still need to be checked against the replies. Next up is a blind human comparison of RLHF and RLJF replies.
Caveats: the experiment covers one task, policies up to 7B, and two to four runs per configuration. The grounded runs used only the 1.5B policy and a short, made-up help center. Skywork's labels come from both people and LLMs (Liu et al., 2025). The reported standard errors measure variation across messages, not between training runs. The length cap was chosen after the uncapped results were known. Menu, hedge and invention counts came from simple regular expressions. The rescoring with the help center was a one-off check.
📚 Further reading
- Illustrating Reinforcement Learning from Human Feedback, the Hugging Face introduction to RLHF.
- Learning to summarize from human feedback (Stiennon et al., 2020) and Training language models to follow instructions with human feedback (Ouyang et al., 2022), the RLHF recipe most later work builds on.
- Constitutional AI: Harmlessness from AI Feedback (Bai et al., 2022) and RLAIF vs. RLHF (Lee et al., 2023), on replacing human labels with model labels.
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Zheng et al., 2023), on the biases of LLM judges.
- Teaching language models to support answers with verified quotes (Menick et al., 2022), which rewards answers backed by quotes from a source.
- Fine-Grained Human Feedback Gives Better Rewards for Language Model Training (Wu et al., 2023), with a separate reward for factual errors.
- Scaling Laws for Reward Model Overoptimization (Gao et al., 2022), on what happens when a policy optimises a proxy reward for too long.
- A Long Way to Go: Investigating Length Correlations in RLHF (Singhal et al., 2023), on how much of RLHF's gain is length.
- DeepSeekMath (Shao et al., 2024), which introduced GRPO.
🙏 Acknowledgements
Huge thanks to the teams behind JEV-9B, Clef-Flash, decider and Laya for the open weights, and to the TRL maintainers for the GRPO trainer.
📌 Citation
@misc{galego2026rljf,
author = {Galego, Jo{\~a}o},
title = {RLJF: Reinforcement Learning from Jev Feedback},
year = {2026},
url = {https://huggingface.co/blog/jgalego/rljf}
}