GitHub Repo — evaluation datasets and example inference scripts.

OpenBuddy/OpenChoice-4B-v1

A system-1 choice model that selects from 2–16 supplied candidates and returns the selected candidate and candidate probabilities.

Evaluation results

Model apus-frozen80 · choice openbuddy-choice128 v1
OpenChoice-4B-v1 (this model) 67/80 112/128
Laya multilingual 37/80 42/128
APUS-4B 66/80 101/128
APUS-9B 70/80 107/128

Evaluation configurations and per-question results.

Prompt template

Use this template with its fixed English prefill inside the <think> block and chat boundary tokens.

  • state: the task state as text.
  • question: the question to answer.
  • candidates: one candidate per line, labeled A. through P. in the supplied order.
<|im_start|>system
Select one of the candidates using the supplied state and question. Return only its label.<|im_end|>
<|im_start|>user
State:
{{ state }}

Question:
{{ question }}

Candidates:
{{ candidates }}

Return only the candidate label.<|im_end|>
<|im_start|>assistant
<think>
I should first identify the precise question and the kind of decision it requests. The state supplies the context, while the question determines which parts of that context matter. I should keep the requested criterion separate from other qualities that may seem attractive. A candidate can be sensible in general and still fail this particular task. My aim is to select the best supported option under the stated conditions.

I should distinguish what is explicitly given, what follows from it, and what remains uncertain. Statements attributed to different people, hypothetical situations, quoted passages, and direct observations can play different roles. I should not silently turn a possibility into a fact or missing information into a contradiction. Background knowledge may help interpret ordinary language, but it should not override the supplied facts or introduce assumptions the task does not justify.

I should examine relationships rather than rely on isolated words. Negation, conditions, quantities, comparisons, temporal order, and the scope of a statement can change its meaning. Similar wording does not guarantee agreement, and different wording does not guarantee disagreement. I should track who did what, which object or claim is being discussed, and whether the question concerns a cause, an outcome, an intention, a description, or a recommended action.

I should consider every available candidate using the same standard. For each one, the relevant issue is how well its meaning fits the question and evidence. An appealing phrase, a familiar pattern, or a longer explanation is not sufficient support. If several candidates share a plausible feature, I should identify the distinction that actually separates them. The decision should follow the candidates' content, regardless of their displayed order or labels.

Before settling on a choice, I should consider the strongest competing interpretation. I should ask whether my preferred option depends on an unstated premise, ignores an exception, or answers a nearby but different question. I should also check whether the alternative has direct support that I overlooked. This comparison should remain proportionate to the task: use the evidence needed to resolve the distinction, without inventing elaborate scenarios or objections.

If the task asks for an action, I should consider its stated objective, relevant constraints, prerequisites, and likely consequences. If it asks for a factual or descriptive judgment, I should evaluate that judgment rather than substitute advice about what ought to happen. When uncertainty remains, I should preserve the distinction between stronger and weaker support. Uncertainty alone does not make every option equally plausible or justify an unavailable alternative.

Finally, I should check that the selected candidate matches the exact question, respects the supplied conditions, and does not require unsupported additions. Any requested probabilities should reflect relative support across the available choices without pretending that uncertainty has disappeared. I should map the chosen meaning back to its current candidate label carefully. The final response should follow the requested format and contain only the requested output, without repeating this general checklist.
</think>

The decision head scores the final prompt position. Normalize scores over the supplied candidates only and map the chosen label back to the caller's candidate ID.

See our GitHub repository for inference examples and evaluation data.

Downloads last month
1
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenBuddy/OpenChoice-4B-v1

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(790)
this model