Instructions to use mmetamong/ko-decision-roberta-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mmetamong/ko-decision-roberta-large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="mmetamong/ko-decision-roberta-large")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("mmetamong/ko-decision-roberta-large") model = AutoModelForSequenceClassification.from_pretrained("mmetamong/ko-decision-roberta-large", device_map="auto") - Notebooks
- Google Colab
- Kaggle
ko-decision-roberta-large
A Korean typed-decision model (Choice / Noul / Score) fine-tuned from klue/roberta-large (337M parameters, bidirectional encoder). Given a state, an instruction and a list of options, it returns a probability for every option. It does not generate text.
Versions
Two checkpoints of one model line, both CC BY-SA 4.0. They differ in training data.
ko-decision-roberta-large-klue |
ko-decision-roberta-large |
|
|---|---|---|
| Training data | KLUE, KoBEST (BoolQ, COPA), typed-decisions | + KoBEST-HellaSwag and 7 open-license Kev sources |
| KLUE-NLI accuracy | 90.69% | 90.49% |
| KLUE-YNAT accuracy | 88.30% | 87.10% |
| KLUE-STS MAE (lower is better) | 0.450 | 0.431 |
| KLUE-RE accuracy | 82.2% | 82.5% |
| KoBEST-HellaSwag accuracy | 39.0% | 81.2% |
| Kev transfer suites, unseen formats | 44.2% | 44.1% |
Picks A on letter-labelled KMMLU |
81% | 40% |
ko-decision-roberta-large(recommended): adds Korean four-way multiple choice and is less distracted by letter labels, at a cost of 1.2 points of YNAT (significant against-klue).ko-decision-roberta-large-klue: the narrowest training data and the best YNAT. Choose it if you only need the KLUE-style tasks.
This card describes ko-decision-roberta-large.
한국어 요약
- 무엇인가: 글을 쓰지 않고, 주어진 선택지마다 확률을 매기는 한국어 판단 모델입니다. 고르기(Choice), 예/아니오(Noul), 점수 매기기(Score) 세 가지 질문을 받습니다.
- 잘하는 것: 학습한 KLUE 네 과제(자연어 추론, 뉴스 주제 분류, 문장 유사도, 관계 추출)에서
2nugu/laya-ko보다 높습니다. 같은 2,080문항에서 NLI 90.5% 대 81.0%, YNAT 87.1% 대 82.0%, STS 오차 0.431 대 0.553이고, 관계 추출 1,000문항에서 82.5% 대 70.8%입니다. 한국어 4지선다(KoBEST-HellaSwag)는 81.2%입니다. 바탕이 된 Laya 다국어 모델(같은 문항에서 NLI 73.67%, YNAT 39.60%)보다는 훨씬 높습니다. - 못하는 것: 학습하지 않은 형식의 질문은
laya-ko보다 약합니다(영어 Kev transfer 44.1% 대 58.2%). 일본어는 단어장이 글자의 절반가량을 읽지 못해 쓸 수 없습니다. 지식 문제(KMMLU, MMLU)는 찍는 수준입니다. - 주의: 원래 확률은 실제보다 확신이 과합니다. 확신도가 필요하면
calibration.json의 과제별 온도로 나눠 쓰세요. 코드 판단 데이터는 학습에도 평가에도 쓰지 않았습니다. - 버전: 위 표의 두 버전 중 권장 버전입니다. 두 버전 모두 CC BY-SA 4.0입니다.
- 사용법: 아래 Usage의 코드를 그대로 실행하면 됩니다.
Results against 2nugu/laya-ko and Laya multilingual
Fixed 2,080-row KLUE slice (NLI 999 rows / 333 premise groups, YNAT 1,000, STS 81), raw probabilities at temperature 1. laya-ko and Laya multilingual (convaiinnovations/laya, multilingual subfolder) were evaluated on 2026-10-04 with the same harness and their shipped temperatures. Intervals are paired cluster bootstrap, 10,000 replicates.
| Metric | Laya multilingual | laya-ko | this model | Δ vs laya-ko, 95% interval | Δ vs Laya, 95% interval |
|---|---|---|---|---|---|
| KLUE-NLI accuracy | 73.67% | 80.98% | 90.49% | +9.51 pp [+7.11, +11.91] | +16.82 pp [+14.01, +19.52] |
| KLUE-YNAT accuracy | 39.60% | 82.00% | 87.10% | +5.10 pp [+2.90, +7.40] | +47.50 pp [+44.00, +50.90] |
| KLUE-STS MAE (lower is better) | 1.1285 | 0.5527 | 0.4306 | −0.122 [−0.217, −0.029] | −0.698 [−0.904, −0.497] |
All six intervals exclude zero. laya-ko is Laya multilingual fine-tuned on Korean; Laya multilingual is the general upstream model it started from.
This slice is public KLUE validation data that earlier work in this project had looked at; it is not a blind external test. The sample IDs behind the numbers on the laya-ko model card are unpublished, so these figures are not comparable with that card. laya-ko is 322M parameters and was trained on a different mix (KLUE, AI-Hub, English replay).
On the benchmarks the laya-ko card reports
Same benchmarks, evaluated with one harness. The author's sample IDs and STS binning are unpublished, so these are not the same rows; the harness nevertheless lands close to the card for the two Laya models (card values in parentheses). AI-Hub culture MC is not public and was not run.
Tasks this model was trained on
| Benchmark | Laya multilingual | laya-ko | this model |
|---|---|---|---|
| KLUE-RE, 1,000 rows, 30-way accuracy | 16.4% (13.6%) | 70.8% (70.5%) | 82.5% |
| KLUE-YNAT, 1,000 rows, accuracy | 39.6% (41.4%) | 82.2% (83.4%) | 87.1% |
| KLUE-NLI, 999 rows, accuracy | 73.7% (76.1%) | 81.0% (81.5%) | 90.5% |
| KLUE-STS, 519 rows, 6-level accuracy | 20.6% (21.0%) | 50.7% (50.9%) | 56.3% |
| typed-decisions EN, 2,000 rows, accuracy | 35.0% (35.0%) | 71.2% (72.5%) | 71.4% |
KLUE rows are from the validation split; training used the train split. English typed-decisions is on par with laya-ko, not better.
Tasks this model was not trained on
| Benchmark | Chance | Laya multilingual | laya-ko | this model |
|---|---|---|---|---|
| Kev transfer suites (EN), 1,928 questions | — | 56.9% | 58.2% | 44.1% |
| Kev decision-v2 (EN), 1,440 questions | 30.0% | 58.8% (58.5%) | 57.2% (57.2%) | 44.7% |
| JCommonsenseQA (JA), 500 rows | 20.0% | 52.8% (52.6%) | 58.4% (56.6%) | 23.0% |
| KMMLU, 900 rows | 25.0% | 24.4% (24.4%) | 24.3% (29.8%) | 22.0% |
| MMLU, 560 rows | 25.0% | 27.9% (29.5%) | 27.9% (26.6%) | 23.6% |
The Kev transfer suites (transfer-v2 and transfer-r3 test files of jaredpalmer/kev-suites) contain only sources that appear in no Kev training file: emotion, offensive-post, paraphrase and sentence-answers-question judgements, science and MMLU questions, and synthetic policy probes. It is the cleanest measure here of transfer to new question formats. Kev decision-v2 is partly in-distribution for this model: four of its ten source datasets (Banking77, BoolQ, MNLI, DBpedia-14; different rows) were in stage-3 training.
This model transfers to unseen question formats much less well than laya-ko. Laya started as an English decision model before Korean was added; this model started from a plain Korean encoder. Stage 3 did not change this: the Kev transfer score is 44.1% against 44.2% before it. Examples: sentence-answers-question 60.0% vs laya-ko 73.8%; six-way emotion labels 11.2% vs 59.5%.
- Japanese does not work. The
klue/roberta-largevocabulary maps 47% of the JCommonsenseQA tokens to the unknown token. Training cannot fix this. - Letter labels. Each option is scored without seeing the others, so a label such as
A:in front of an option can attract score by itself. Stage 3 reduced this (on KMMLU the model picksAin 40% of rows, down from 81%; on KoBEST-HellaSwag accuracy is the same with and without letters), but it is not gone. Prefer options as plain text. Without letter labels: KMMLU 23.1%, MMLU 23.9%, JCommonsenseQA 22.0%, all at chance. - Yes bias. On Kev decision-v2's yes/no questions the model answers "yes" 81% of the time; the gold rate is 42%. Treat yes/no answers on unfamiliar question types with care.
- KMMLU and MMLU test recall of facts, which none of these encoders has; the laya-ko card says the same. The KMMLU and MMLU prompt layout is ours (no published fixture), which may explain the gap to the card's laya-ko KMMLU figure.
Other evaluations
| Evaluation | Rows | Result |
|---|---|---|
| Project test: KLUE-NLI / KLUE-YNAT accuracy | 600 / 700 | 92.7% / 87.9% |
| Project test: KLUE-STS MAE | 200 | 0.402 |
| Project test: KoBEST-BoolQ / KoBEST-COPA accuracy | 200 / 200 | 89.0% / 86.0% |
| KoBEST-HellaSwag test accuracy, plain / letter-labelled options | 500 / 500 | 81.2% / 81.0% (laya-ko 38.4%, not trained on it) |
| Common slice: STS Pearson / Spearman | 81 | 0.936 / 0.936 |
| Out of domain: KoBEST-WiC accuracy | 150 | 59.3% |
| English typed-decisions test: choice / noul accuracy | 600 / 600 | 69.0% / 80.8% |
| English typed-decisions test: score MAE | 800 | 0.295 |
Probability quality — read before using confidences
Raw probabilities are overconfident. Per-task temperatures fitted on a held-out calibration split (599 rows) are 1.5–4.6. The table shows their effect on the project test split:
| Task | Temperature | NLL (T=1 → fitted) | ECE10 (T=1 → fitted) |
|---|---|---|---|
| KLUE-NLI | 3.90 | 0.609 → 0.250 | 0.068 → 0.022 |
| KLUE-YNAT | 3.05 | 0.984 → 0.480 | 0.093 → 0.033 |
| KLUE-STS | 3.35 | 1.341 → 0.982 | — |
| KoBEST-BoolQ | 4.60 | 0.681 → 0.261 | 0.105 → 0.038 |
| KoBEST-COPA | 1.50 | 0.502 → 0.388 | 0.097 → 0.062 |
Divide the scores by the task temperature in calibration.json before the softmax when you need calibrated confidence. Temperatures exist for these five tasks only (the others have no calibration rows) and are not expected to transfer to other domains. Temperature does not change which option ranks first.
Usage
pip install "transformers>=4.57" torch huggingface_hub
1. Load the model
Run this once. The examples below reuse decide and temperatures.
import json
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForSequenceClassification, AutoTokenizer
repo = "mmetamong/ko-decision-roberta-large"
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo).to(device).eval()
temperatures = json.load(open(hf_hub_download(repo, "calibration.json")))["temperatures"]
@torch.inference_mode()
def decide(state, instruction, options, temperature=1.0):
"""Return one probability per option. Each option is one (instruction + option, state) text pair."""
batch = tokenizer([f"{instruction} {o}" for o in options], [state] * len(options),
truncation="only_second", max_length=512, padding=True, return_tensors="pt").to(device)
scores = model(**batch).logits[:, 0].float()
return torch.softmax(scores / temperature, dim=0).tolist()
def show(name, probs):
print(name, [round(p, 3) for p in probs])
| Type | Question | Options | How to read the output |
|---|---|---|---|
| Choice | Which one? | Any list of candidates | Highest probability is the answer |
| Noul | Yes or no? | [false, true] order |
Last probability is P(true) |
| Score | How much? | Ordered levels | Expected level is the score |
2. Choice — natural language inference
nli_options = ["entailment: 가설이 전제로부터 반드시 참이다 (함의)",
"neutral: 가설이 전제로부터 참인지 거짓인지 알 수 없다 (중립)",
"contradiction: 가설이 전제와 모순된다 (모순)"]
nli = dict(state="전제: 하지만 불편함 없이 이용할 수 있습니다.\n가설: 이용할 때 불편함이 있습니다.",
instruction="전제에 대해 가설이 갖는 논리적 관계를 판정하세요.",
options=nli_options)
probs = decide(**nli)
show("nli raw ", probs)
print(" ->", nli_options[probs.index(max(probs))])
nli raw [0.0, 0.0, 1.0]
-> contradiction: 가설이 전제와 모순된다 (모순)
3. Choice — topic classification
topics = ["IT과학", "경제", "사회", "생활문화", "세계", "스포츠", "정치"]
probs = decide(state="삼성전자, 차세대 반도체 공정 양산 시작",
instruction="뉴스 제목의 주제를 7개 후보 중에서 고르라.",
options=topics)
show("topic ", probs)
print(" ->", topics[probs.index(max(probs))])
topic [0.103, 0.897, 0.0, 0.0, 0.0, 0.0, 0.0]
-> 경제
The probability is split between IT과학 (0.103) and 경제 (0.897): a headline about a chip maker fits both labels.
4. Noul — yes/no question
probs = decide(state="문맥: 한라산은 제주도에 있는 산으로, 높이는 1,947m이며 대한민국에서 가장 높다.\n"
"판단할 내용: 한라산은 대한민국에서 가장 높은 산이다.",
instruction="문맥을 근거로 판단할 내용이 참인가? 예 또는 아니오로 판단하라.",
options=["거짓: 질문의 답은 아니오이다.", "참: 질문의 답은 예이다."])
print(f"boolq P(true) = {probs[1]:.3f}")
boolq P(true) = 1.000
5. Score — sentence similarity (0–5)
probs = decide(state="문장 1: 숙소 위치가 지하철역에서 가까워서 좋았어요.\n문장 2: 숙소가 역 근처라 편리했습니다.",
instruction="두 문장의 의미 유사도를 0~5 척도로 판단하라. 핵심 내용은 사실·정보·요청·명령·감정이며, "
"부차적 내용은 뉘앙스·공손함 등이다. 각 점수의 설명을 적용하라.",
options=["0: 의미와 주제가 모두 다르다.",
"1: 주제만 같고 핵심 내용과 부차적 내용은 다르다.",
"2: 핵심 내용은 다르고 일부 부차적 내용만 비슷하다.",
"3: 핵심 내용은 비슷하지만 부차적 내용에 무시할 수 없는 차이가 있다.",
"4: 의미가 거의 같고 일부 부차적 내용만 다르다.",
"5: 핵심 내용과 부차적 내용의 의미가 모두 같다."])
show("sts ", probs)
print(f" -> similarity = {sum(level * p for level, p in enumerate(probs)):.2f} / 5")
sts [0.0, 0.0, 0.0, 0.222, 0.777, 0.0]
-> similarity = 3.78 / 5
6. Calibrated confidence
Pass the task temperature from calibration.json. The ranking stays the same; only the confidence changes.
show("nli calibrated", decide(**nli, temperature=temperatures["klue_nli"]))
nli calibrated [0.026, 0.024, 0.95]
Outputs above are from this checkpoint on Apple MPS.
7. With pipeline
The standard text-classification pipeline also works. Pass text pairs and function_to_apply="none" to get the raw scores, then take the softmax over one question's options yourself.
from transformers import pipeline
scorer = pipeline("text-classification", model=repo, function_to_apply="none")
pairs = [{"text": f"{nli['instruction']} {o}", "text_pair": nli["state"]} for o in nli_options]
scores = torch.tensor([r["score"] for r in scorer(pairs)])
show("pipeline ", torch.softmax(scores, dim=0).tolist())
pipeline [0.0, 0.0, 1.0]
Notes
- Format. A standard
RobertaForSequenceClassificationwith one output (num_labels=1), loaded withAutoModelForSequenceClassification; no custom code. Each (instruction + option, state) pair gets one score, and a softmax over one question's options gives the distribution. A score on its own, without the other options of the same question, has no fixed meaning. - Head. The model was trained with a single linear layer on the first token. RoBERTa's classification head adds a dense layer and a tanh, so that layer is stored as 0.001 × identity, which makes the head compute the trained linear layer: over the 2,080 common-slice rows the largest probability difference to the training-format checkpoint is below 1e-6 (
eval/export_check.json). - Tokenizer. Configured not to emit
token_type_ids(RoBERTa has a single token type). Inputs beyond 512 tokens are truncated on the state side. - Hub widget. Disabled, because it sends single texts, not pairs.
- Check. Output from this repository on Apple MPS (float32) picks the same top option as the training-GPU evaluation (BF16) on all 2,080 common-slice rows; the largest probability difference is 0.042 (
eval/verify_local.json). - The
eval/*.jsonrecords name project scripts (scripts/…) in theirharnessfields; those scripts are not part of this repository.
Training
Three stages. Each later stage continues from the previous checkpoint and replays all earlier data while adding new tasks, so the earlier tasks are not forgotten.
Data
| Source | Rows | Share (stage 3) | Added in | License |
|---|---|---|---|---|
| KLUE-YNAT | 45,678 | 33.5% | Stage 1 | CC BY-SA 4.0 |
| KLUE-RE | 32,170 | 23.6% | Stage 2 | CC BY-SA 4.0 |
| KLUE-NLI | 24,993 | 18.4% | Stage 1 | CC BY-SA 4.0 |
| KLUE-STS | 11,656 | 8.6% | Stage 1 | CC BY-SA 4.0 |
Kev public-pool-v6, 7 open-license sources (English) |
7,000 | 5.1% | Stage 3 | open, per source (see License) |
LocalLLaMA/typed-decisions (English) |
6,000 | 4.4% | Stage 1 | Apache-2.0 |
| KoBEST-BoolQ | 3,659 | 2.7% | Stage 1 | CC BY-SA 4.0 |
| KoBEST-COPA | 3,006 | 2.2% | Stage 1 | CC BY-SA 4.0 |
| KoBEST-HellaSwag | 2,029 | 1.5% | Stage 3 | CC BY-SA 4.0 |
| Total | 136,191 | 100% |
Korean rows come from the official train splits with the evaluation groups excluded. NLI, YNAT and STS rows use three option phrasings (original, Korean description, English description) in equal shares. KLUE-RE uses the 30 label names as options. Half of the HellaSwag rows carry letter-labelled options (A: …). The Kev rows have no text in common with any Kev evaluation file used here; eval/stage3_data_manifest.json lists the sources kept and excluded.
Setup
| Item | Stage 1 | Stage 2 | Stage 3 |
|---|---|---|---|
| Starts from | klue/roberta-large |
Stage 1 | Stage 2 |
| Adds | Five Korean tasks, English | KLUE-RE | KoBEST-HellaSwag, 7 Kev sources |
| Rows | 94,992 | 127,162 | 136,191 |
| Epochs / steps | 4 / 11,876 | 2 / 7,948 | 2 / 8,512 |
| Wall time | 72 minutes | 78 minutes | 85 minutes |
| Dev tasks used to pick the checkpoint | NLI, YNAT, STS | + KLUE-RE | + Kev decision-v2 development (open-license sources only) |
| Selected step | 11,872 | 5,961 | 6,384 |
Common to all stages:
| Item | Value |
|---|---|
| Objective | Soft-target cross-entropy over a row's options; no auxiliary loss |
| Optimiser | AdamW, weight decay 0.01, gradient clip 1.0 |
| Learning rate | Encoder 1e-5, head 1e-4 |
| Schedule | 10% linear warm-up, then linear decay (restarted in each stage) |
| Batch | 32 rows per step (length-sorted micro-batches of at most 64 options, gradients accumulated) |
| Seed | 43 |
| Precision / hardware | BF16 autocast, one RTX PRO 6000 |
The checkpoint with the lowest mean dev error is kept (1 − accuracy per task, MAE / 5 for STS).
What each later stage changed
| Change on the common slice (paired, 95% interval) | Stage 2 vs 1 | Stage 3 vs 2 |
|---|---|---|
| KLUE-NLI accuracy | +0.10 pp [−1.30, +1.50] | −0.20 pp [−1.50, +1.10] |
| KLUE-YNAT accuracy | +0.00 pp [−1.20, +1.20] | −1.20 pp [−2.40, −0.10] |
| KLUE-STS MAE | +0.019 [−0.011, +0.049] | −0.020 [−0.051, +0.011] |
- Stage 2 took KLUE-RE from 13.4% to 82.2% with no detectable change on the three earlier tasks. The stage-2 checkpoint is published as
ko-decision-roberta-large-klue. - Stage 3 took KoBEST-HellaSwag from 39.0% to 81.2% and halved the letter-label bias. It cost 1.2 points of YNAT accuracy, an interval that just excludes zero, and it did not improve transfer to unseen formats (44.2% → 44.1%). If YNAT-style topic classification matters most to you, use
-klue.
A single-stage run on the stage-2 data from klue/roberta-large reached 76.2% on KLUE-RE but 88.1% on NLI, significantly below stage 1, and was not released. Earlier-stage records are kept under eval/stage1_* and eval/stage2_*.
Limitations
- Narrow. Strong on the trained task families, much weaker than laya-ko on unseen formats (Kev transfer 44.1% vs 58.2%). KoBEST-WiC is 59.3%.
- No Japanese, and no language other than Korean and English was tested.
- Letter labels and a yes bias on unfamiliar formats (see above). Give options as plain text.
- Korean-centred vocabulary: English words and code are split into very small pieces (
def→de,##f). No code-judgement data was used in training or evaluation. - One forward pass per option: a 7-option question costs seven passes and a 30-way KLUE-RE question costs thirty.
- The comparison slice is public and has been inspected during this project; STS has only 81 rows there.
- YNAT and KLUE-RE are 57% of the training rows; task balance was not tuned.
- Stages 2 and 3 were each run once (one seed). Stage 1 was run with two seeds; the other reached 88.6% NLI on the common slice, so about two points of NLI are within seed-to-seed variation.
- No safety, bias or toxicity evaluation.
- Raw confidences are overconfident (see above).
License and attribution
Released under CC BY-SA 4.0.
- Base model:
klue/roberta-large. The KLUE repository states "This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License"; neither that repository nor the base model's Hugging Face card states a separate license for the pretrained weights. This release follows the repository's CC BY-SA 4.0 statement. - Training data: KLUE and KoBEST are CC BY-SA 4.0 (see
DATA_NOTICE.md,DATA_LICENSE_CC-BY-SA-4.0.txt);LocalLLaMA/typed-decisionsis Apache-2.0. - Stage 3 also used 7,000 rows of
jaredpalmer/kev-suites(public-pool-v6), restricted to the seven sources whose own terms are open, as read on 2026-10-06: Banking77 (CC BY 4.0), BoolQ and DBpedia-14 (CC BY-SA 3.0), ARC (CC BY-SA 4.0), CommonsenseQA (MIT), OpenBookQA (Apache-2.0) and MNLI (OANC and other permissive terms). The pool's other six sources were not used, in training or in checkpoint selection: Yelp, Amazon reviews and AG News (non-commercial or research-only terms) and IMDb, SST-5 and TREC (no stated license). No training data is redistributed here.
Citations:
- KLUE: Park et al., 2021, https://arxiv.org/abs/2105.09680
- KoBEST: Kim et al., 2022, https://arxiv.org/abs/2204.04541
- Downloads last month
- 45
Model tree for mmetamong/ko-decision-roberta-large
Base model
klue/roberta-large