Text Generation
Transformers
Safetensors
English
qwen3
medical
clinical
biomedical
instruction-following
tool-calling
function-calling
from-scratch
KOS-V5
university-of-kentucky
university-of-louisville
ccts
conversational
text-generation-inference
Instructions to use Kentucky-Open-Science/KOS-V5-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Kentucky-Open-Science/KOS-V5-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Kentucky-Open-Science/KOS-V5-Instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Kentucky-Open-Science/KOS-V5-Instruct") model = AutoModelForCausalLM.from_pretrained("Kentucky-Open-Science/KOS-V5-Instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Kentucky-Open-Science/KOS-V5-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Kentucky-Open-Science/KOS-V5-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kentucky-Open-Science/KOS-V5-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Kentucky-Open-Science/KOS-V5-Instruct
- SGLang
How to use Kentucky-Open-Science/KOS-V5-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Kentucky-Open-Science/KOS-V5-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kentucky-Open-Science/KOS-V5-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Kentucky-Open-Science/KOS-V5-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kentucky-Open-Science/KOS-V5-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Kentucky-Open-Science/KOS-V5-Instruct with Docker Model Runner:
docker model run hf.co/Kentucky-Open-Science/KOS-V5-Instruct
File size: 22,193 Bytes
41f25fa cea427e 41f25fa cea427e 41f25fa cea427e f61a416 cea427e c554a92 cea427e c554a92 cea427e e8d1a9b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 | ---
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
language:
- en
developed_by: University of Kentucky (College of Medicine Office for Research; Center for Clinical and Translational Science) and University of Louisville (Kentucky Center for Digital Innovation)
affiliations:
- name: University of Kentucky, College of Medicine Office for Research
url: https://medicine.uky.edu/sites/research
- name: University of Kentucky, Center for Clinical and Translational Science (CCTS)
url: https://www.ccts.uky.edu/
- name: University of Louisville, Kentucky Center for Digital Innovation
url: https://centers.louisville.edu/kentucky-center-digital-innovation
model_name: KOS-V5-Instruct
model_codename: Catbird
model_type: qwen3
base_model: Kentucky-Open-Science/KOS-V5-Base
tags:
- medical
- clinical
- biomedical
- instruction-following
- tool-calling
- function-calling
- from-scratch
- qwen3
- KOS-V5
- university-of-kentucky
- university-of-louisville
- ccts
---
<p align="center">
<img src="catbird_llm_logo.png" alt="Catbird" width="320"/>
</p>
# KOS-V5-Instruct · *"Catbird"*
**Developed by**
**University of Kentucky**
- [College of Medicine, Office for Research](https://medicine.uky.edu/sites/research)
- [Center for Clinical and Translational Science (CCTS)](https://www.ccts.uky.edu/)
**University of Louisville**
- [Kentucky Center for Digital Innovation](https://centers.louisville.edu/kentucky-center-digital-innovation)
**A 3.72B-parameter medical language model trained from scratch.** It is not distilled, not pruned and not
continued-pretrained from a general base. KOS-V5 (codename **Catbird**) is the fifth-generation Kentucky Open
Science model line. This repository holds the **instruction-tuned head** of that line: the
[KOS-V5-Base](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) pretraining checkpoint, taken through
SFT and two GRPO reinforcement-learning legs.
Unlike the base, this model **follows instructions and calls tools**. It is the downstream SFT/RL artifact that
[KOS-V5-Base](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) was built to initialise.
**Code name: Catbird.** The KOS-V5 series is nicknamed Catbird; native to Kentucky, the Gray Catbird
(*Dumetella carolinensis*) is a medium-sized songbird famous for its distinct, cat-like "meow" call. This LLM
was trained completely from scratch by teams from the University of Kentucky (Cat) and University of Louisville
(Bird), so the code name is fitting.
> ⚠️ **Research use only.** This model is provided for research purposes only and must not be used for any
> commercial, clinical, legal, or production-grade application. The user assumes all risks associated with its use.
---
Its instruction ability comes from **GRPO reinforcement learning** against the *official* IFEval verifier, and its
tool-calling ability from a **second GRPO leg** against the *official* BFCL AST checker, on a base that ranks
**first of 17** at modelling held-out clinical text.
**IFEval reported as strict-avg** = `(prompt-level strict + instruction-level strict) / 2` — the exact metric the
Hugging Face Open LLM Leaderboard publishes as "IFEval."
| IFEval **strict-avg** | model | who built it, and how |
| --: | :-- | :-- |
| **72.19** | **KOS-V5-Instruct (ours)** | University research team, 235B tokens, from scratch |
| 64.7 | Qwen2.5-3B-Instruct | Alibaba, ~18 trillion tokens |
| **61.6** | [KOS-V4-Instruct](https://huggingface.co/Kentucky-Open-Science/KOS-V4-Instruct) (previous generation) | University research team, 180B tokens, 24 GPUs |
| 55.9 | GPT-3.5-turbo-1106 (the original ChatGPT) | OpenAI, ~10,000-GPU supercomputer |
KOS-V5-Instruct **improves on KOS-V4-Instruct across every benchmark measured**: IFEval strict-avg
**61.6 → 72.19** (+10.6), MMLU **0.2782 → 0.4512** (+17.3), medical QA (**PubMedQA 0.7060**, MedQA 0.3802,
MedMCQA 0.3648 — all up on V4), and official BFCL function-calling
**72.75/73.00/60.50 → 85.00/84.00/80.50** (+12.3 / +11.0 / +20.0). It clears the original GPT-3.5-turbo
generation and the commercially trained Qwen2.5-3B on instruction following, and its tool calling now runs
**above the Qwen3-4B-Instruct-2507 peer**.
## Core specifications
| Attribute | Detail |
| :--- | :--- |
| **Architecture** | Decoder-only Transformer (`Qwen3ForCausalLM`), Grouped-Query Attention |
| **Parameters** | 3.715 B |
| **Hidden / Layers** | 2560 / 36 |
| **Attention** | 32 query / 8 KV heads (GQA 4:1), head_dim 128, per-head QK-RMSNorm |
| **Feed-forward** | SwiGLU, intermediate 9728 |
| **Vocabulary** | 32,000, custom medical byte-level BPE |
| **Context length** | 32,768 |
| **Position encoding** | RoPE, θ = 25,000 |
| **Embeddings** | tied |
| **Precision** | bfloat16 (7.43 GB, single shard) |
## Pre-training (the KOS-V5 base)
Fine-tuned from [**KOS-V5-Base**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) — the from-scratch
pretrained foundation, a complete single-epoch run over **235.2B tokens**. See that card for corpus composition and
disclosed pretraining issues.
## Post-training (this model)
Three stages on top of the base — no LoRA, no distillation, no reward model, no LLM judge.
**SFT** — one shuffled full-parameter pass over a **736,990-record / 1.32B-token** audited instruction mix
(ChatML). The mix was deduplicated, instruction-collision canonicalized, structurally validated, per-record
decontaminated and BFCL-closure scanned; clinical FHIR records were dropped and tool-record system prompts
de-welded so tool use binds to the request rather than to a fixed frame.
**RL leg 1 — instruction following (GRPO via verl)** — deterministic verifiable reward. The reward is the
**official** `lm_eval` IFEval instruction registry — the same checkers the benchmark scores with, not a
re-implementation. One 8-GPU wall, 164 steps, KL 0.001 (`low_var_kl`), rollout n=16.
**RL leg 2 — tool calling (GRPO via verl)** — a second leg seeded from leg 1. The reward is the **official**
BFCL `ast_checker` (`bfcl_eval`). Each prompt renders its tool schemas through the model's **own** chat template
(byte-exact to the official `tools=` rendering), and the prompt set is filtered to only rows the official checker
can grade. 8-GPU wall, KL 0.001, rollout n=16; this repository ships the **step-40** checkpoint, selected for the
best BFCL / abstention balance and least policy drift. BFCL rose **76.8/71.0/69.5 → 85.0/84.0/80.5** with
instruction following, grounded abstention and knowledge all held.
**Forgetting control** — out-of-distribution broad-holdout perplexity at **0.99× the pre-RL base** (8.88 vs 8.97),
measured on a web crawl postdating the training corpus. No measurable forgetting.
## The medical foundation
**This is a medical model.** KOS-V5-Instruct inherits a base trained on a **54-source medical/biomedical corpus**
— not a general-purpose model with medical fine-tuning bolted on.
The strongest evidence is **bits-per-byte on held-out medical text**, which is tokenizer-agnostic and therefore
the only strictly fair cross-model comparison. In a **17-model pool** — including dedicated biomedical
specialists BioMedLM (300B PubMed tokens), Meditron-7B, PMC-LLaMA-7B and MedGemma-4B — the KOS-V5 base ranks
**first**:
| medical text (BPB, lower is better) | KOS-V5-Base | rank |
| :-- | --: | --: |
| **5-corpus mean, held-out medical text** | **0.4635** | **1 / 17** |
| clinical narratives | **0.4179** | **1 / 17** |
| radiology | **0.5132** | **1 / 17** |
| chest X-ray reports | **0.6688** | **1 / 17** |
| BIOSSES biomedical sentence similarity (Pearson / Spearman) | **0.7097 / 0.7014** | **1 / 17** |
| BLURB biomedical probe mean | 0.7268 | 2 / 17 |
Every comparator in that pool was trained on **1.3–153× more data** (0.3–36T tokens vs our 0.235T). See
[KOS-V5-Base](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) for the full 96-metric evaluation.
### Medical MMLU (from the 57-subject run above)
The 9 medical subjects of MMLU, extracted from the same official 5-shot run:
| medical subject | KOS-V5-Instruct | KOS-V4-Instruct |
| :-- | --: | --: |
| high-school biology | 0.5774 | 0.2387 |
| clinical knowledge | 0.5623 | 0.3170 |
| nutrition | 0.5359 | 0.2843 |
| college biology | 0.5347 | 0.2917 |
| medical genetics | 0.5100 | 0.2700 |
| anatomy | 0.4815 | 0.3185 |
| professional medicine | 0.4375 | 0.2132 |
| college medicine | 0.4046 | 0.2486 |
| virology | 0.3795 | 0.2952 |
| **medical-9 mean** | **0.4915** | **0.2752** |
**+21.6 points over KOS-V4-Instruct**, and above the model's own full-MMLU average (0.4512) — the medical
subjects are where it is strongest.
### Medical QA benchmarks (official suites)
Official `medqa_4options`, `medmcqa` and `pubmedqa` from the same pristine lm-evaluation-harness,
5-shot, loglikelihood, metric `acc`.
| medical benchmark | **KOS-V5-Instruct** | KOS-V4-Instruct | YuLan-Mini-Instruct | marin-8b-instruct | Qwen3-4B-Instruct-2507 |
| :-- | --: | --: | --: | --: | --: |
| params | **3.7B** | 3.0B | 2.4B | 8.0B | 4.0B |
| **PubMedQA** | **0.7060** | 0.6860 | 0.6960 | 0.7500 | 0.7720 |
| MedQA (USMLE, 4-option) | **0.3802** | 0.2820 | 0.3511 | 0.4878 | 0.6159 |
| MedMCQA | **0.3648** | 0.2778 | 0.3856 | 0.4961 | 0.5804 |
**KOS-V5-Instruct improves on KOS-V4-Instruct on all three** (+9.8 MedQA, +8.7 MedMCQA, +2.0 PubMedQA).
**PubMedQA is the standout: 0.7060**, ahead of YuLan-Mini and within reach of Stanford's Marin-8B at **less than
half the parameters**. PubMedQA tests comprehension of **biomedical literature** — the closest of these three to
what the base was actually trained on. The USMLE-style exam MCQs (MedQA, MedMCQA) are where the answer-letter
bottleneck below bites hardest.
> ⚠️ **Why the MCQ numbers understate this model.** Our own measurements show KOS models place very little
> probability mass on MCQ answer *letters*: the format, not the knowledge, is the bottleneck. A model that ranks
> **1 of 17** at modelling clinical text while scoring modestly on multiple-choice is exhibiting exactly that gap.
> **Read the BPB results as the medical signal and the MCQ results as a floor, not a ceiling.**
## Evaluation & benchmarks
**Official suites only**, EleutherAI lm-evaluation-harness `0.4.12.dev0` at upstream commit `c1c4bea`, run from a
**pristine clone** with stock, unmodified task definitions.
- **IFEval** — stock `ifeval` task, 0-shot, greedy (`do_sample=false`, `temperature=0.0`),
task-default `max_gen_toks=1280`, `--apply_chat_template`, seed 0. Constraint checking by the harness's vendored
**Google** verifier (`instructions_registry`, 25 instruction types).
- **MMLU** — stock `mmlu` group, official 57 subjects / 14,042 test items, 5-shot from `dev` (`first_n`),
loglikelihood over A–D, metric `acc` (not `acc_norm`), **no chat template**.
- **Medical QA** — stock `medqa_4options`, `medmcqa`, `pubmedqa` tasks, 5-shot, loglikelihood, metric `acc`,
no chat template.
| benchmark | KOS-V5-Instruct | KOS-V4-Instruct | Δ |
| :-- | --: | --: | --: |
| **IFEval strict-avg** | **72.19** | 61.6 | **+10.6** |
| IFEval prompt-strict | 0.6728 | 0.5471 | +0.126 |
| IFEval inst-strict | 0.7710 | 0.6655 | +0.106 |
| IFEval prompt-loose | 0.6932 | 0.5693 | +0.124 |
| IFEval inst-loose | 0.7878 | 0.6882 | +0.100 |
| **MMLU (57-subj, 5-shot, `acc`)** | **0.4512** | 0.2782 | **+0.173** |
### Tool / function calling — official BFCL
Measured with the **official `bfcl_eval`** suite in **FC (function-calling) mode**, non-live categories, the model
prompted in its own native tool format and served via vLLM. **Tool calling is a trained objective of this model**
— the second GRPO leg optimised the official BFCL AST checker directly.
| BFCL (official, FC mode, **non-live AST**) | **KOS-V5-Instruct** | Qwen3-4B-Instruct-2507 (peer) | KOS-V4-Instruct |
| :-- | --: | --: | --: |
| simple *(334/400)* | **85.00** | 83.20 | 72.75 |
| multiple *(157/200)* | **84.00** | 79.00 | 73.00 |
| parallel *(147/200)* | **80.50** | 73.50 | 60.50 |
**KOS-V5-Instruct is above the Qwen3-4B-Instruct-2507 peer on all three BFCL categories**, and far above the
previous KOS-V4-Instruct. This is the axis the tool-calling GRPO leg was built to move, and it moved.
> **Scope.** These are the **non-live AST** categories only (`simple_python`, `multiple`, `parallel`). The
> live, multi-turn, web-search and memory categories were **not run**, so no BFCL *overall* score is reported
> here — the suite's aggregate column is not meaningful when most categories are unrun.
> **Engine note.** These BFCL numbers come from the official `bfcl_eval` harness on a **vLLM** backend, whereas
> the IFEval and MMLU figures on this card come from the HuggingFace backend of a pristine lm-evaluation-harness.
> Both are official suites; they are not the same inference stack, and that is stated rather than blurred.
**Cross-harness reproduction.** IFEval strict-avg measured **72.19** (pristine HF harness) and **72.0** (our
RL-evaluation harness) in two independent runs — a **0.19-point** agreement across two harness builds, far below
the benchmark's own ±2.14-point standard error on 541 prompts, so they are the same measurement.
**Harness validation.** The identical pipeline scored the peer mark **Qwen3-4B-Instruct-2507 at 84.71** IFEval
strict-avg on the same pristine harness, and independently reproduced KOS-V4-Instruct's MMLU to four decimal
places (0.2782). A score of 0.0 on this pipeline would therefore be a model property, not a harness failure.
### Grounded abstention & robustness — official RGB
Measured on the **official RGB harness** (retrieval-augmented generation benchmark).
| RGB (official) | KOS-V5-Instruct | Qwen3-4B-Instruct-2507 (peer) |
| :-- | --: | --: |
| **negative rejection** (declines the unanswerable) | **57.33** | 39.0 |
| noise robustness | 64.0 | 93.67 |
**Grounded abstention is a genuine strength: neg-reject 57.33 vs the peer's 39.0** — this model declines to answer
unanswerable questions far more often than it invents an answer. Noise-robustness (64.0) improved over an earlier
revision (58.67) but remains below the 70 threshold we treat as a pass.
### IFEval in context (strict-avg)
| model | weights | company | params | IFEval strict-avg |
| :-- | :-- | :-- | :-- | --: |
| GPT-4o-mini | Proprietary | OpenAI | 8B + | 79 \* |
| Llama-3.2-3B-Instruct | Open | Meta | 3.2B | 73.9 |
| **KOS-V5-Instruct (ours)** | **Open** | **Univ. of Kentucky / Louisville** | **3.7B** | **72.19** |
| Qwen2.5-3B-Instruct | Open | Alibaba | 3.0B | 64.7 |
| Phi-3-medium-4k-instruct | Open | Microsoft | 14.0B | 64.2 |
| Mistral-Large | Proprietary | Mistral AI | 46.7B + | 63 \* |
| **KOS-V4-Instruct (previous gen)** | **Open** | **Univ. of Kentucky** | **3.0B** | **61.6** |
| Yi-1.5-9B-Chat | Open | 01.AI | 8.8B | 60.5 |
| Phi-3.5-mini-instruct | Open | Microsoft | 3.8B | 57.7 |
| GPT-3.5-turbo-0613 | Proprietary | OpenAI | 20B + | 57 \* |
| Phi-3-mini-4k-instruct | Open | Microsoft | 3.8B | 56.1 |
| GPT-3.5-turbo-1106 | Proprietary | OpenAI | 20B + | 55.9 |
| Mistral-7B-Instruct-v0.2 | Open | Mistral AI | 7.2B | 55.0 |
| Llama-3.1-8B-Instruct | Open | Meta | 8.0B | 44.3 |
| Llama-2-13b-chat | Open | Meta | 13.0B | 39.8 |
**\* strict estimate** — no official IFEval strict sub-metrics published; estimated from published AVG4 or
prompt-strict (loose metrics run ~2–4 pts above strict). **\+ unofficial params.**
### Against university-built instruction models
Measured by us on the identical pristine harness, same protocol:
| model | institution | params | IFEval strict-avg | MMLU |
| :-- | :-- | --: | --: | --: |
| **KOS-V5-Instruct (ours)** | **UK / UofL** | **3.7B** | **72.19** | **0.4512** |
| marin-8b-instruct | Stanford | 8.0B | 70.83 | 0.6112 |
| YuLan-Mini-Instruct | Renmin | 2.4B | 61.51 | 0.5278 |
| **KOS-V4-Instruct (ours)** | **UK** | **3.0B** | **60.63** | 0.2782 |
| LLäMmlein-7B-chat | Würzburg | 7.0B | 54.07 | 0.5252 |
| Poro-34B-chat | U Turku | 34.2B | 34.63 | — |
| Minerva-7B-instruct | Sapienza | 7.4B | 21.51 | 0.4071 |
| CroissantLLMChat | CentraleSupélec | 1.3B | 19.94 | 0.2401 |
| Tucano-2b4-Instruct | U Bonn | 2.4B | 14.95 | 0.2589 |
On instruction following KOS-V5-Instruct now places **first among nine university-built instruct models**, ahead of
Stanford's Marin-8B (70.83) at less than half its parameters, and of Poro-34B at **9× its parameter count**. Note
that several of these models are non-English-first (Finnish, Italian, French, Portuguese, German) and are being
measured on English benchmarks, which understates their designed capability. On parametric knowledge (MMLU) the
larger, more heavily trained models still lead.
## Retrieval & embeddings
Beyond generation, KOS-V5-Instruct **also serves as a dense text retriever.** A companion **LoRA adapter** —
[**KOS-V5-Retriever**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Retriever) — converts
this model into an embedding model (llm2vec-style: bidirectional attention + mean-pooling + a contrastively-trained
rank-32 LoRA), with these base weights **frozen and unchanged**.
On the **official BEIR SciFact** benchmark (the `beir` library + pytrec_eval — the public-leaderboard scorer),
**zero-shot** (training excluded SciFact, verified clean), it scores **NDCG@10 = 0.7007** (Recall@10 0.864) — a
**strong dense retriever**, above BM25 (0.665) and within the GTR/E5/BGE band (0.70–0.76). The adapter is
hot-swappable: attach it for retrieval, detach it for generation. See the
[adapter card](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Retriever) for the encode recipe and
the full retrieval details.
## Known regressions and limitations
> These are disclosed deliberately. A high benchmark score does not make this checkpoint production-ready.
- **Tool / function calling is a trained strength — and now measured against the peer.** The second GRPO leg
optimised the *official* BFCL AST checker directly; BFCL is 85.0/84.0/80.5, above the peer and far above the
previous KOS generation. Earlier internal KOS tool-calling figures are deliberately omitted: several were
measured against benchmark data the model had been trained on and are recorded in our own audit as **invalid**.
- **Grounded abstention / fabrication IS measured — and strong.** On the official RGB harness, negative-rejection
is 57.33 (peer 39.0). *Earlier revisions of this card stated fabrication was NOT independently measured; it now
is.* Noise-robustness (64.0) is still below our 70 pass bar, so retrieval-noise handling remains a known gap.
- **The peer leads on knowledge and raw instruction following.** IFEval 72.19 vs the peer's 84.71; MMLU 0.4512 vs
0.7266; medical QA below the peer. This is a from-scratch 3.7B model on 235B tokens against one trained on orders
of magnitude more data — strong for its scale, not state-of-the-art in absolute terms.
- **MMLU 0.4512 is above chance (0.25) but modest.** This is not a knowledge model; it should not be used as a
medical question-answering authority.
- **Not a medical-MCQ model.** As with KOS-V4, do not benchmark or deploy it as one; read the BPB results (base
card) as the medical signal and the MCQ results as a floor.
## Data contamination
- **IFEval: CLEAN (verbatim).** The SFT mix and the IFEval RL prompt set were exact-containment scanned against
IFEval's official 541 test prompts — **0 exact containments**. Exact matching cannot detect paraphrase or
reformatting.
- **Tool calling (BFCL): CLEAN (verbatim).** The tool-calling RL prompt pool was exact-containment scanned against
**5,437 full-length official BFCL prompts — 0 exact containments**.
- **MMLU / PubMedQA / MedQA / MedMCQA: UNCHECKED.** Contamination against these four has **not** been scanned for
this checkpoint. Those numbers should be read with that caveat.
## Prompt / chat format (ChatML)
```
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
Write a haiku about Kentucky. Do not use any commas.<|im_end|>
<|im_start|>assistant
```
Tool / function calling uses the model's native `<tools>` … `</tools>` schema block and `<tool_call>` … `</tool_call>`
response format; pass your function schemas via the tokenizer's `apply_chat_template(..., tools=[...])`.
## Quickstart
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Kentucky-Open-Science/KOS-V5-Instruct"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
messages = [{"role": "user", "content": "Write a haiku about Kentucky. Do not use any commas."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
```
## Deployment
| Precision | Approx. VRAM | Notes |
| :--- | :--- | :--- |
| **bfloat16** | ~9 GB | native weights (7.43 GB) + activations; a single 16 GB GPU is comfortable |
## Related models
- [**KOS-V5-Retriever**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Retriever) — a LoRA
**retrieval/embedding adapter** for this model (llm2vec; official BEIR SciFact NDCG@10 **0.70**, zero-shot).
- [**KOS-V5-Base**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) — the from-scratch pretrained
foundation this model is tuned from (3.72B, 235.2B tokens).
- [**KOS-V4-Instruct**](https://huggingface.co/Kentucky-Open-Science/KOS-V4-Instruct) — previous generation
(3.0B, IFEval 61.6), the public release.
- [**KOS-V4-Base**](https://huggingface.co/Kentucky-Open-Science/KOS-V4-Base) — previous-generation foundation
(3.015B, 180.3B tokens).
## Intended use & limitations
Research use only. English only. Not for clinical, commercial, legal, or production-grade use. Outputs may be
factually wrong or fabricated. This model must not be used to make or inform medical decisions.
## Naming
The program is **KOS** (KOS-V1..V6). Earlier internal names are not used.
|