ko-decision-roberta-large

A Korean typed-decision model (Choice / Noul / Score) fine-tuned from klue/roberta-large (337M parameters, bidirectional encoder). Given a state, an instruction and a list of options, it returns a probability for every option. It does not generate text.

Versions

Two checkpoints of one model line, both CC BY-SA 4.0. They differ in training data.

ko-decision-roberta-large-klue ko-decision-roberta-large
Training data KLUE, KoBEST (BoolQ, COPA), typed-decisions + KoBEST-HellaSwag and 7 open-license Kev sources
KLUE-NLI accuracy 90.69% 90.49%
KLUE-YNAT accuracy 88.30% 87.10%
KLUE-STS MAE (lower is better) 0.450 0.431
KLUE-RE accuracy 82.2% 82.5%
KoBEST-HellaSwag accuracy 39.0% 81.2%
Kev transfer suites, unseen formats 44.2% 44.1%
Picks A on letter-labelled KMMLU 81% 40%
  • ko-decision-roberta-large (recommended): adds Korean four-way multiple choice and is less distracted by letter labels, at a cost of 1.2 points of YNAT (significant against -klue).
  • ko-decision-roberta-large-klue: the narrowest training data and the best YNAT. Choose it if you only need the KLUE-style tasks.

This card describes ko-decision-roberta-large.

한국어 요약

  • 무엇인가: 글을 쓰지 않고, 주어진 선택지마다 확률을 매기는 한국어 판단 모델입니다. 고르기(Choice), 예/아니오(Noul), 점수 매기기(Score) 세 가지 질문을 받습니다.
  • 잘하는 것: 학습한 KLUE 네 과제(자연어 추론, 뉴스 주제 분류, 문장 유사도, 관계 추출)에서 2nugu/laya-ko보다 높습니다. 같은 2,080문항에서 NLI 90.5% 대 81.0%, YNAT 87.1% 대 82.0%, STS 오차 0.431 대 0.553이고, 관계 추출 1,000문항에서 82.5% 대 70.8%입니다. 한국어 4지선다(KoBEST-HellaSwag)는 81.2%입니다. 바탕이 된 Laya 다국어 모델(같은 문항에서 NLI 73.67%, YNAT 39.60%)보다는 훨씬 높습니다.
  • 못하는 것: 학습하지 않은 형식의 질문은 laya-ko보다 약합니다(영어 Kev transfer 44.1% 대 58.2%). 일본어는 단어장이 글자의 절반가량을 읽지 못해 쓸 수 없습니다. 지식 문제(KMMLU, MMLU)는 찍는 수준입니다.
  • 주의: 원래 확률은 실제보다 확신이 과합니다. 확신도가 필요하면 calibration.json의 과제별 온도로 나눠 쓰세요. 코드 판단 데이터는 학습에도 평가에도 쓰지 않았습니다.
  • 버전: 위 표의 두 버전 중 권장 버전입니다. 두 버전 모두 CC BY-SA 4.0입니다.
  • 사용법: 아래 Usage의 코드를 그대로 실행하면 됩니다.

Results against 2nugu/laya-ko and Laya multilingual

Fixed 2,080-row KLUE slice (NLI 999 rows / 333 premise groups, YNAT 1,000, STS 81), raw probabilities at temperature 1. laya-ko and Laya multilingual (convaiinnovations/laya, multilingual subfolder) were evaluated on 2026-10-04 with the same harness and their shipped temperatures. Intervals are paired cluster bootstrap, 10,000 replicates.

Metric Laya multilingual laya-ko this model Δ vs laya-ko, 95% interval Δ vs Laya, 95% interval
KLUE-NLI accuracy 73.67% 80.98% 90.49% +9.51 pp [+7.11, +11.91] +16.82 pp [+14.01, +19.52]
KLUE-YNAT accuracy 39.60% 82.00% 87.10% +5.10 pp [+2.90, +7.40] +47.50 pp [+44.00, +50.90]
KLUE-STS MAE (lower is better) 1.1285 0.5527 0.4306 −0.122 [−0.217, −0.029] −0.698 [−0.904, −0.497]

All six intervals exclude zero. laya-ko is Laya multilingual fine-tuned on Korean; Laya multilingual is the general upstream model it started from.

This slice is public KLUE validation data that earlier work in this project had looked at; it is not a blind external test. The sample IDs behind the numbers on the laya-ko model card are unpublished, so these figures are not comparable with that card. laya-ko is 322M parameters and was trained on a different mix (KLUE, AI-Hub, English replay).

On the benchmarks the laya-ko card reports

Same benchmarks, evaluated with one harness. The author's sample IDs and STS binning are unpublished, so these are not the same rows; the harness nevertheless lands close to the card for the two Laya models (card values in parentheses). AI-Hub culture MC is not public and was not run.

Tasks this model was trained on

Benchmark Laya multilingual laya-ko this model
KLUE-RE, 1,000 rows, 30-way accuracy 16.4% (13.6%) 70.8% (70.5%) 82.5%
KLUE-YNAT, 1,000 rows, accuracy 39.6% (41.4%) 82.2% (83.4%) 87.1%
KLUE-NLI, 999 rows, accuracy 73.7% (76.1%) 81.0% (81.5%) 90.5%
KLUE-STS, 519 rows, 6-level accuracy 20.6% (21.0%) 50.7% (50.9%) 56.3%
typed-decisions EN, 2,000 rows, accuracy 35.0% (35.0%) 71.2% (72.5%) 71.4%

KLUE rows are from the validation split; training used the train split. English typed-decisions is on par with laya-ko, not better.

Tasks this model was not trained on

Benchmark Chance Laya multilingual laya-ko this model
Kev transfer suites (EN), 1,928 questions — 56.9% 58.2% 44.1%
Kev decision-v2 (EN), 1,440 questions 30.0% 58.8% (58.5%) 57.2% (57.2%) 44.7%
JCommonsenseQA (JA), 500 rows 20.0% 52.8% (52.6%) 58.4% (56.6%) 23.0%
KMMLU, 900 rows 25.0% 24.4% (24.4%) 24.3% (29.8%) 22.0%
MMLU, 560 rows 25.0% 27.9% (29.5%) 27.9% (26.6%) 23.6%

The Kev transfer suites (transfer-v2 and transfer-r3 test files of jaredpalmer/kev-suites) contain only sources that appear in no Kev training file: emotion, offensive-post, paraphrase and sentence-answers-question judgements, science and MMLU questions, and synthetic policy probes. It is the cleanest measure here of transfer to new question formats. Kev decision-v2 is partly in-distribution for this model: four of its ten source datasets (Banking77, BoolQ, MNLI, DBpedia-14; different rows) were in stage-3 training.

This model transfers to unseen question formats much less well than laya-ko. Laya started as an English decision model before Korean was added; this model started from a plain Korean encoder. Stage 3 did not change this: the Kev transfer score is 44.1% against 44.2% before it. Examples: sentence-answers-question 60.0% vs laya-ko 73.8%; six-way emotion labels 11.2% vs 59.5%.

  • Japanese does not work. The klue/roberta-large vocabulary maps 47% of the JCommonsenseQA tokens to the unknown token. Training cannot fix this.
  • Letter labels. Each option is scored without seeing the others, so a label such as A: in front of an option can attract score by itself. Stage 3 reduced this (on KMMLU the model picks A in 40% of rows, down from 81%; on KoBEST-HellaSwag accuracy is the same with and without letters), but it is not gone. Prefer options as plain text. Without letter labels: KMMLU 23.1%, MMLU 23.9%, JCommonsenseQA 22.0%, all at chance.
  • Yes bias. On Kev decision-v2's yes/no questions the model answers "yes" 81% of the time; the gold rate is 42%. Treat yes/no answers on unfamiliar question types with care.
  • KMMLU and MMLU test recall of facts, which none of these encoders has; the laya-ko card says the same. The KMMLU and MMLU prompt layout is ours (no published fixture), which may explain the gap to the card's laya-ko KMMLU figure.

Other evaluations

Evaluation Rows Result
Project test: KLUE-NLI / KLUE-YNAT accuracy 600 / 700 92.7% / 87.9%
Project test: KLUE-STS MAE 200 0.402
Project test: KoBEST-BoolQ / KoBEST-COPA accuracy 200 / 200 89.0% / 86.0%
KoBEST-HellaSwag test accuracy, plain / letter-labelled options 500 / 500 81.2% / 81.0% (laya-ko 38.4%, not trained on it)
Common slice: STS Pearson / Spearman 81 0.936 / 0.936
Out of domain: KoBEST-WiC accuracy 150 59.3%
English typed-decisions test: choice / noul accuracy 600 / 600 69.0% / 80.8%
English typed-decisions test: score MAE 800 0.295

Probability quality — read before using confidences

Raw probabilities are overconfident. Per-task temperatures fitted on a held-out calibration split (599 rows) are 1.5–4.6. The table shows their effect on the project test split:

Task Temperature NLL (T=1 → fitted) ECE10 (T=1 → fitted)
KLUE-NLI 3.90 0.609 → 0.250 0.068 → 0.022
KLUE-YNAT 3.05 0.984 → 0.480 0.093 → 0.033
KLUE-STS 3.35 1.341 → 0.982 —
KoBEST-BoolQ 4.60 0.681 → 0.261 0.105 → 0.038
KoBEST-COPA 1.50 0.502 → 0.388 0.097 → 0.062

Divide the scores by the task temperature in calibration.json before the softmax when you need calibrated confidence. Temperatures exist for these five tasks only (the others have no calibration rows) and are not expected to transfer to other domains. Temperature does not change which option ranks first.

Usage

pip install "transformers>=4.57" torch huggingface_hub

1. Load the model

Run this once. The examples below reuse decide and temperatures.

import json
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForSequenceClassification, AutoTokenizer

repo = "mmetamong/ko-decision-roberta-large"
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"

tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo).to(device).eval()
temperatures = json.load(open(hf_hub_download(repo, "calibration.json")))["temperatures"]


@torch.inference_mode()
def decide(state, instruction, options, temperature=1.0):
    """Return one probability per option. Each option is one (instruction + option, state) text pair."""
    batch = tokenizer([f"{instruction} {o}" for o in options], [state] * len(options),
                      truncation="only_second", max_length=512, padding=True, return_tensors="pt").to(device)
    scores = model(**batch).logits[:, 0].float()
    return torch.softmax(scores / temperature, dim=0).tolist()


def show(name, probs):
    print(name, [round(p, 3) for p in probs])
Type Question Options How to read the output
Choice Which one? Any list of candidates Highest probability is the answer
Noul Yes or no? [false, true] order Last probability is P(true)
Score How much? Ordered levels Expected level is the score

2. Choice — natural language inference

nli_options = ["entailment: 가설이 전제로부터 반드시 참이다 (함의)",
               "neutral: 가설이 전제로부터 참인지 거짓인지 알 수 없다 (중립)",
               "contradiction: 가설이 전제와 모순된다 (모순)"]
nli = dict(state="전제: 하지만 불편함 없이 이용할 수 있습니다.\n가설: 이용할 때 불편함이 있습니다.",
           instruction="전제에 대해 가설이 갖는 논리적 관계를 판정하세요.",
           options=nli_options)
probs = decide(**nli)
show("nli raw       ", probs)
print("  ->", nli_options[probs.index(max(probs))])
nli raw        [0.0, 0.0, 1.0]
  -> contradiction: 가설이 전제와 모순된다 (모순)

3. Choice — topic classification

topics = ["IT과학", "경제", "사회", "생활문화", "세계", "스포츠", "정치"]
probs = decide(state="삼성전자, 차세대 반도체 공정 양산 시작",
               instruction="뉴스 제목의 주제를 7개 후보 중에서 고르라.",
               options=topics)
show("topic         ", probs)
print("  ->", topics[probs.index(max(probs))])
topic          [0.103, 0.897, 0.0, 0.0, 0.0, 0.0, 0.0]
  -> 경제

The probability is split between IT과학 (0.103) and 경제 (0.897): a headline about a chip maker fits both labels.

4. Noul — yes/no question

probs = decide(state="문맥: 한라산은 제주도에 있는 산으로, 높이는 1,947m이며 대한민국에서 가장 높다.\n"
                     "판단할 내용: 한라산은 대한민국에서 가장 높은 산이다.",
               instruction="문맥을 근거로 판단할 내용이 참인가? 예 또는 아니오로 판단하라.",
               options=["거짓: 질문의 답은 아니오이다.", "참: 질문의 답은 예이다."])
print(f"boolq           P(true) = {probs[1]:.3f}")
boolq           P(true) = 1.000

5. Score — sentence similarity (0–5)

probs = decide(state="문장 1: 숙소 위치가 지하철역에서 가까워서 좋았어요.\n문장 2: 숙소가 역 근처라 편리했습니다.",
               instruction="두 문장의 의미 유사도를 0~5 척도로 판단하라. 핵심 내용은 사실·정보·요청·명령·감정이며, "
                           "부차적 내용은 뉘앙스·공손함 등이다. 각 점수의 설명을 적용하라.",
               options=["0: 의미와 주제가 모두 다르다.",
                        "1: 주제만 같고 핵심 내용과 부차적 내용은 다르다.",
                        "2: 핵심 내용은 다르고 일부 부차적 내용만 비슷하다.",
                        "3: 핵심 내용은 비슷하지만 부차적 내용에 무시할 수 없는 차이가 있다.",
                        "4: 의미가 거의 같고 일부 부차적 내용만 다르다.",
                        "5: 핵심 내용과 부차적 내용의 의미가 모두 같다."])
show("sts           ", probs)
print(f"  -> similarity = {sum(level * p for level, p in enumerate(probs)):.2f} / 5")
sts            [0.0, 0.0, 0.0, 0.222, 0.777, 0.0]
  -> similarity = 3.78 / 5

6. Calibrated confidence

Pass the task temperature from calibration.json. The ranking stays the same; only the confidence changes.

show("nli calibrated", decide(**nli, temperature=temperatures["klue_nli"]))
nli calibrated [0.026, 0.024, 0.95]

Outputs above are from this checkpoint on Apple MPS.

7. With pipeline

The standard text-classification pipeline also works. Pass text pairs and function_to_apply="none" to get the raw scores, then take the softmax over one question's options yourself.

from transformers import pipeline

scorer = pipeline("text-classification", model=repo, function_to_apply="none")
pairs = [{"text": f"{nli['instruction']} {o}", "text_pair": nli["state"]} for o in nli_options]
scores = torch.tensor([r["score"] for r in scorer(pairs)])
show("pipeline      ", torch.softmax(scores, dim=0).tolist())
pipeline       [0.0, 0.0, 1.0]

Notes

  • Format. A standard RobertaForSequenceClassification with one output (num_labels=1), loaded with AutoModelForSequenceClassification; no custom code. Each (instruction + option, state) pair gets one score, and a softmax over one question's options gives the distribution. A score on its own, without the other options of the same question, has no fixed meaning.
  • Head. The model was trained with a single linear layer on the first token. RoBERTa's classification head adds a dense layer and a tanh, so that layer is stored as 0.001 × identity, which makes the head compute the trained linear layer: over the 2,080 common-slice rows the largest probability difference to the training-format checkpoint is below 1e-6 (eval/export_check.json).
  • Tokenizer. Configured not to emit token_type_ids (RoBERTa has a single token type). Inputs beyond 512 tokens are truncated on the state side.
  • Hub widget. Disabled, because it sends single texts, not pairs.
  • Check. Output from this repository on Apple MPS (float32) picks the same top option as the training-GPU evaluation (BF16) on all 2,080 common-slice rows; the largest probability difference is 0.042 (eval/verify_local.json).
  • The eval/*.json records name project scripts (scripts/…) in their harness fields; those scripts are not part of this repository.

Training

Three stages. Each later stage continues from the previous checkpoint and replays all earlier data while adding new tasks, so the earlier tasks are not forgotten.

Data

Source Rows Share (stage 3) Added in License
KLUE-YNAT 45,678 33.5% Stage 1 CC BY-SA 4.0
KLUE-RE 32,170 23.6% Stage 2 CC BY-SA 4.0
KLUE-NLI 24,993 18.4% Stage 1 CC BY-SA 4.0
KLUE-STS 11,656 8.6% Stage 1 CC BY-SA 4.0
Kev public-pool-v6, 7 open-license sources (English) 7,000 5.1% Stage 3 open, per source (see License)
LocalLLaMA/typed-decisions (English) 6,000 4.4% Stage 1 Apache-2.0
KoBEST-BoolQ 3,659 2.7% Stage 1 CC BY-SA 4.0
KoBEST-COPA 3,006 2.2% Stage 1 CC BY-SA 4.0
KoBEST-HellaSwag 2,029 1.5% Stage 3 CC BY-SA 4.0
Total 136,191 100%

Korean rows come from the official train splits with the evaluation groups excluded. NLI, YNAT and STS rows use three option phrasings (original, Korean description, English description) in equal shares. KLUE-RE uses the 30 label names as options. Half of the HellaSwag rows carry letter-labelled options (A: …). The Kev rows have no text in common with any Kev evaluation file used here; eval/stage3_data_manifest.json lists the sources kept and excluded.

Setup

Item Stage 1 Stage 2 Stage 3
Starts from klue/roberta-large Stage 1 Stage 2
Adds Five Korean tasks, English KLUE-RE KoBEST-HellaSwag, 7 Kev sources
Rows 94,992 127,162 136,191
Epochs / steps 4 / 11,876 2 / 7,948 2 / 8,512
Wall time 72 minutes 78 minutes 85 minutes
Dev tasks used to pick the checkpoint NLI, YNAT, STS + KLUE-RE + Kev decision-v2 development (open-license sources only)
Selected step 11,872 5,961 6,384

Common to all stages:

Item Value
Objective Soft-target cross-entropy over a row's options; no auxiliary loss
Optimiser AdamW, weight decay 0.01, gradient clip 1.0
Learning rate Encoder 1e-5, head 1e-4
Schedule 10% linear warm-up, then linear decay (restarted in each stage)
Batch 32 rows per step (length-sorted micro-batches of at most 64 options, gradients accumulated)
Seed 43
Precision / hardware BF16 autocast, one RTX PRO 6000

The checkpoint with the lowest mean dev error is kept (1 − accuracy per task, MAE / 5 for STS).

What each later stage changed

Change on the common slice (paired, 95% interval) Stage 2 vs 1 Stage 3 vs 2
KLUE-NLI accuracy +0.10 pp [−1.30, +1.50] −0.20 pp [−1.50, +1.10]
KLUE-YNAT accuracy +0.00 pp [−1.20, +1.20] −1.20 pp [−2.40, −0.10]
KLUE-STS MAE +0.019 [−0.011, +0.049] −0.020 [−0.051, +0.011]
  • Stage 2 took KLUE-RE from 13.4% to 82.2% with no detectable change on the three earlier tasks. The stage-2 checkpoint is published as ko-decision-roberta-large-klue.
  • Stage 3 took KoBEST-HellaSwag from 39.0% to 81.2% and halved the letter-label bias. It cost 1.2 points of YNAT accuracy, an interval that just excludes zero, and it did not improve transfer to unseen formats (44.2% → 44.1%). If YNAT-style topic classification matters most to you, use -klue.

A single-stage run on the stage-2 data from klue/roberta-large reached 76.2% on KLUE-RE but 88.1% on NLI, significantly below stage 1, and was not released. Earlier-stage records are kept under eval/stage1_* and eval/stage2_*.

Limitations

  • Narrow. Strong on the trained task families, much weaker than laya-ko on unseen formats (Kev transfer 44.1% vs 58.2%). KoBEST-WiC is 59.3%.
  • No Japanese, and no language other than Korean and English was tested.
  • Letter labels and a yes bias on unfamiliar formats (see above). Give options as plain text.
  • Korean-centred vocabulary: English words and code are split into very small pieces (def → de, ##f). No code-judgement data was used in training or evaluation.
  • One forward pass per option: a 7-option question costs seven passes and a 30-way KLUE-RE question costs thirty.
  • The comparison slice is public and has been inspected during this project; STS has only 81 rows there.
  • YNAT and KLUE-RE are 57% of the training rows; task balance was not tuned.
  • Stages 2 and 3 were each run once (one seed). Stage 1 was run with two seeds; the other reached 88.6% NLI on the common slice, so about two points of NLI are within seed-to-seed variation.
  • No safety, bias or toxicity evaluation.
  • Raw confidences are overconfident (see above).

License and attribution

Released under CC BY-SA 4.0.

  • Base model: klue/roberta-large. The KLUE repository states "This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License"; neither that repository nor the base model's Hugging Face card states a separate license for the pretrained weights. This release follows the repository's CC BY-SA 4.0 statement.
  • Training data: KLUE and KoBEST are CC BY-SA 4.0 (see DATA_NOTICE.md, DATA_LICENSE_CC-BY-SA-4.0.txt); LocalLLaMA/typed-decisions is Apache-2.0.
  • Stage 3 also used 7,000 rows of jaredpalmer/kev-suites (public-pool-v6), restricted to the seven sources whose own terms are open, as read on 2026-10-06: Banking77 (CC BY 4.0), BoolQ and DBpedia-14 (CC BY-SA 3.0), ARC (CC BY-SA 4.0), CommonsenseQA (MIT), OpenBookQA (Apache-2.0) and MNLI (OANC and other permissive terms). The pool's other six sources were not used, in training or in checkpoint selection: Yelp, Amazon reviews and AG News (non-commercial or research-only terms) and IMDb, SST-5 and TREC (no stated license). No training data is redistributed here.

Citations:

Downloads last month
45
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mmetamong/ko-decision-roberta-large

Finetuned
(80)
this model

Datasets used to train mmetamong/ko-decision-roberta-large

Papers for mmetamong/ko-decision-roberta-large