--- license: apache-2.0 library_name: transformers pipeline_tag: text-generation language: - en developed_by: University of Kentucky (College of Medicine Office for Research; Center for Clinical and Translational Science) and University of Louisville (Kentucky Center for Digital Innovation) affiliations: - name: University of Kentucky, College of Medicine Office for Research url: https://medicine.uky.edu/sites/research - name: University of Kentucky, Center for Clinical and Translational Science (CCTS) url: https://www.ccts.uky.edu/ - name: University of Louisville, Kentucky Center for Digital Innovation url: https://centers.louisville.edu/kentucky-center-digital-innovation model_name: KOS-V5-Instruct model_codename: Catbird model_type: qwen3 base_model: Kentucky-Open-Science/KOS-V5-Base tags: - medical - clinical - biomedical - instruction-following - tool-calling - function-calling - from-scratch - qwen3 - KOS-V5 - university-of-kentucky - university-of-louisville - ccts ---

Catbird

# KOS-V5-Instruct · *"Catbird"* **Developed by** **University of Kentucky** - [College of Medicine, Office for Research](https://medicine.uky.edu/sites/research) - [Center for Clinical and Translational Science (CCTS)](https://www.ccts.uky.edu/) **University of Louisville** - [Kentucky Center for Digital Innovation](https://centers.louisville.edu/kentucky-center-digital-innovation) **A 3.72B-parameter medical language model trained from scratch.** It is not distilled, not pruned and not continued-pretrained from a general base. KOS-V5 (codename **Catbird**) is the fifth-generation Kentucky Open Science model line. This repository holds the **instruction-tuned head** of that line: the [KOS-V5-Base](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) pretraining checkpoint, taken through SFT and two GRPO reinforcement-learning legs. Unlike the base, this model **follows instructions and calls tools**. It is the downstream SFT/RL artifact that [KOS-V5-Base](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) was built to initialise. **Code name: Catbird.** The KOS-V5 series is nicknamed Catbird; native to Kentucky, the Gray Catbird (*Dumetella carolinensis*) is a medium-sized songbird famous for its distinct, cat-like "meow" call. This LLM was trained completely from scratch by teams from the University of Kentucky (Cat) and University of Louisville (Bird), so the code name is fitting. > ⚠️ **Research use only.** This model is provided for research purposes only and must not be used for any > commercial, clinical, legal, or production-grade application. The user assumes all risks associated with its use. --- Its instruction ability comes from **GRPO reinforcement learning** against the *official* IFEval verifier, and its tool-calling ability from a **second GRPO leg** against the *official* BFCL AST checker, on a base that ranks **first of 17** at modelling held-out clinical text. **IFEval reported as strict-avg** = `(prompt-level strict + instruction-level strict) / 2` — the exact metric the Hugging Face Open LLM Leaderboard publishes as "IFEval." | IFEval **strict-avg** | model | who built it, and how | | --: | :-- | :-- | | **72.19** | **KOS-V5-Instruct (ours)** | University research team, 235B tokens, from scratch | | 64.7 | Qwen2.5-3B-Instruct | Alibaba, ~18 trillion tokens | | **61.6** | [KOS-V4-Instruct](https://huggingface.co/Kentucky-Open-Science/KOS-V4-Instruct) (previous generation) | University research team, 180B tokens, 24 GPUs | | 55.9 | GPT-3.5-turbo-1106 (the original ChatGPT) | OpenAI, ~10,000-GPU supercomputer | KOS-V5-Instruct **improves on KOS-V4-Instruct across every benchmark measured**: IFEval strict-avg **61.6 → 72.19** (+10.6), MMLU **0.2782 → 0.4512** (+17.3), medical QA (**PubMedQA 0.7060**, MedQA 0.3802, MedMCQA 0.3648 — all up on V4), and official BFCL function-calling **72.75/73.00/60.50 → 85.00/84.00/80.50** (+12.3 / +11.0 / +20.0). It clears the original GPT-3.5-turbo generation and the commercially trained Qwen2.5-3B on instruction following, and its tool calling now runs **above the Qwen3-4B-Instruct-2507 peer**. ## Core specifications | Attribute | Detail | | :--- | :--- | | **Architecture** | Decoder-only Transformer (`Qwen3ForCausalLM`), Grouped-Query Attention | | **Parameters** | 3.715 B | | **Hidden / Layers** | 2560 / 36 | | **Attention** | 32 query / 8 KV heads (GQA 4:1), head_dim 128, per-head QK-RMSNorm | | **Feed-forward** | SwiGLU, intermediate 9728 | | **Vocabulary** | 32,000, custom medical byte-level BPE | | **Context length** | 32,768 | | **Position encoding** | RoPE, θ = 25,000 | | **Embeddings** | tied | | **Precision** | bfloat16 (7.43 GB, single shard) | ## Pre-training (the KOS-V5 base) Fine-tuned from [**KOS-V5-Base**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) — the from-scratch pretrained foundation, a complete single-epoch run over **235.2B tokens**. See that card for corpus composition and disclosed pretraining issues. ## Post-training (this model) Three stages on top of the base — no LoRA, no distillation, no reward model, no LLM judge. **SFT** — one shuffled full-parameter pass over a **736,990-record / 1.32B-token** audited instruction mix (ChatML). The mix was deduplicated, instruction-collision canonicalized, structurally validated, per-record decontaminated and BFCL-closure scanned; clinical FHIR records were dropped and tool-record system prompts de-welded so tool use binds to the request rather than to a fixed frame. **RL leg 1 — instruction following (GRPO via verl)** — deterministic verifiable reward. The reward is the **official** `lm_eval` IFEval instruction registry — the same checkers the benchmark scores with, not a re-implementation. One 8-GPU wall, 164 steps, KL 0.001 (`low_var_kl`), rollout n=16. **RL leg 2 — tool calling (GRPO via verl)** — a second leg seeded from leg 1. The reward is the **official** BFCL `ast_checker` (`bfcl_eval`). Each prompt renders its tool schemas through the model's **own** chat template (byte-exact to the official `tools=` rendering), and the prompt set is filtered to only rows the official checker can grade. 8-GPU wall, KL 0.001, rollout n=16; this repository ships the **step-40** checkpoint, selected for the best BFCL / abstention balance and least policy drift. BFCL rose **76.8/71.0/69.5 → 85.0/84.0/80.5** with instruction following, grounded abstention and knowledge all held. **Forgetting control** — out-of-distribution broad-holdout perplexity at **0.99× the pre-RL base** (8.88 vs 8.97), measured on a web crawl postdating the training corpus. No measurable forgetting. ## The medical foundation **This is a medical model.** KOS-V5-Instruct inherits a base trained on a **54-source medical/biomedical corpus** — not a general-purpose model with medical fine-tuning bolted on. The strongest evidence is **bits-per-byte on held-out medical text**, which is tokenizer-agnostic and therefore the only strictly fair cross-model comparison. In a **17-model pool** — including dedicated biomedical specialists BioMedLM (300B PubMed tokens), Meditron-7B, PMC-LLaMA-7B and MedGemma-4B — the KOS-V5 base ranks **first**: | medical text (BPB, lower is better) | KOS-V5-Base | rank | | :-- | --: | --: | | **5-corpus mean, held-out medical text** | **0.4635** | **1 / 17** | | clinical narratives | **0.4179** | **1 / 17** | | radiology | **0.5132** | **1 / 17** | | chest X-ray reports | **0.6688** | **1 / 17** | | BIOSSES biomedical sentence similarity (Pearson / Spearman) | **0.7097 / 0.7014** | **1 / 17** | | BLURB biomedical probe mean | 0.7268 | 2 / 17 | Every comparator in that pool was trained on **1.3–153× more data** (0.3–36T tokens vs our 0.235T). See [KOS-V5-Base](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) for the full 96-metric evaluation. ### Medical MMLU (from the 57-subject run above) The 9 medical subjects of MMLU, extracted from the same official 5-shot run: | medical subject | KOS-V5-Instruct | KOS-V4-Instruct | | :-- | --: | --: | | high-school biology | 0.5774 | 0.2387 | | clinical knowledge | 0.5623 | 0.3170 | | nutrition | 0.5359 | 0.2843 | | college biology | 0.5347 | 0.2917 | | medical genetics | 0.5100 | 0.2700 | | anatomy | 0.4815 | 0.3185 | | professional medicine | 0.4375 | 0.2132 | | college medicine | 0.4046 | 0.2486 | | virology | 0.3795 | 0.2952 | | **medical-9 mean** | **0.4915** | **0.2752** | **+21.6 points over KOS-V4-Instruct**, and above the model's own full-MMLU average (0.4512) — the medical subjects are where it is strongest. ### Medical QA benchmarks (official suites) Official `medqa_4options`, `medmcqa` and `pubmedqa` from the same pristine lm-evaluation-harness, 5-shot, loglikelihood, metric `acc`. | medical benchmark | **KOS-V5-Instruct** | KOS-V4-Instruct | YuLan-Mini-Instruct | marin-8b-instruct | Qwen3-4B-Instruct-2507 | | :-- | --: | --: | --: | --: | --: | | params | **3.7B** | 3.0B | 2.4B | 8.0B | 4.0B | | **PubMedQA** | **0.7060** | 0.6860 | 0.6960 | 0.7500 | 0.7720 | | MedQA (USMLE, 4-option) | **0.3802** | 0.2820 | 0.3511 | 0.4878 | 0.6159 | | MedMCQA | **0.3648** | 0.2778 | 0.3856 | 0.4961 | 0.5804 | **KOS-V5-Instruct improves on KOS-V4-Instruct on all three** (+9.8 MedQA, +8.7 MedMCQA, +2.0 PubMedQA). **PubMedQA is the standout: 0.7060**, ahead of YuLan-Mini and within reach of Stanford's Marin-8B at **less than half the parameters**. PubMedQA tests comprehension of **biomedical literature** — the closest of these three to what the base was actually trained on. The USMLE-style exam MCQs (MedQA, MedMCQA) are where the answer-letter bottleneck below bites hardest. > ⚠️ **Why the MCQ numbers understate this model.** Our own measurements show KOS models place very little > probability mass on MCQ answer *letters*: the format, not the knowledge, is the bottleneck. A model that ranks > **1 of 17** at modelling clinical text while scoring modestly on multiple-choice is exhibiting exactly that gap. > **Read the BPB results as the medical signal and the MCQ results as a floor, not a ceiling.** ## Evaluation & benchmarks **Official suites only**, EleutherAI lm-evaluation-harness `0.4.12.dev0` at upstream commit `c1c4bea`, run from a **pristine clone** with stock, unmodified task definitions. - **IFEval** — stock `ifeval` task, 0-shot, greedy (`do_sample=false`, `temperature=0.0`), task-default `max_gen_toks=1280`, `--apply_chat_template`, seed 0. Constraint checking by the harness's vendored **Google** verifier (`instructions_registry`, 25 instruction types). - **MMLU** — stock `mmlu` group, official 57 subjects / 14,042 test items, 5-shot from `dev` (`first_n`), loglikelihood over A–D, metric `acc` (not `acc_norm`), **no chat template**. - **Medical QA** — stock `medqa_4options`, `medmcqa`, `pubmedqa` tasks, 5-shot, loglikelihood, metric `acc`, no chat template. | benchmark | KOS-V5-Instruct | KOS-V4-Instruct | Δ | | :-- | --: | --: | --: | | **IFEval strict-avg** | **72.19** | 61.6 | **+10.6** | | IFEval prompt-strict | 0.6728 | 0.5471 | +0.126 | | IFEval inst-strict | 0.7710 | 0.6655 | +0.106 | | IFEval prompt-loose | 0.6932 | 0.5693 | +0.124 | | IFEval inst-loose | 0.7878 | 0.6882 | +0.100 | | **MMLU (57-subj, 5-shot, `acc`)** | **0.4512** | 0.2782 | **+0.173** | ### Tool / function calling — official BFCL Measured with the **official `bfcl_eval`** suite in **FC (function-calling) mode**, non-live categories, the model prompted in its own native tool format and served via vLLM. **Tool calling is a trained objective of this model** — the second GRPO leg optimised the official BFCL AST checker directly. | BFCL (official, FC mode, **non-live AST**) | **KOS-V5-Instruct** | Qwen3-4B-Instruct-2507 (peer) | KOS-V4-Instruct | | :-- | --: | --: | --: | | simple *(334/400)* | **85.00** | 83.20 | 72.75 | | multiple *(157/200)* | **84.00** | 79.00 | 73.00 | | parallel *(147/200)* | **80.50** | 73.50 | 60.50 | **KOS-V5-Instruct is above the Qwen3-4B-Instruct-2507 peer on all three BFCL categories**, and far above the previous KOS-V4-Instruct. This is the axis the tool-calling GRPO leg was built to move, and it moved. > **Scope.** These are the **non-live AST** categories only (`simple_python`, `multiple`, `parallel`). The > live, multi-turn, web-search and memory categories were **not run**, so no BFCL *overall* score is reported > here — the suite's aggregate column is not meaningful when most categories are unrun. > **Engine note.** These BFCL numbers come from the official `bfcl_eval` harness on a **vLLM** backend, whereas > the IFEval and MMLU figures on this card come from the HuggingFace backend of a pristine lm-evaluation-harness. > Both are official suites; they are not the same inference stack, and that is stated rather than blurred. **Cross-harness reproduction.** IFEval strict-avg measured **72.19** (pristine HF harness) and **72.0** (our RL-evaluation harness) in two independent runs — a **0.19-point** agreement across two harness builds, far below the benchmark's own ±2.14-point standard error on 541 prompts, so they are the same measurement. **Harness validation.** The identical pipeline scored the peer mark **Qwen3-4B-Instruct-2507 at 84.71** IFEval strict-avg on the same pristine harness, and independently reproduced KOS-V4-Instruct's MMLU to four decimal places (0.2782). A score of 0.0 on this pipeline would therefore be a model property, not a harness failure. ### Grounded abstention & robustness — official RGB Measured on the **official RGB harness** (retrieval-augmented generation benchmark). | RGB (official) | KOS-V5-Instruct | Qwen3-4B-Instruct-2507 (peer) | | :-- | --: | --: | | **negative rejection** (declines the unanswerable) | **57.33** | 39.0 | | noise robustness | 64.0 | 93.67 | **Grounded abstention is a genuine strength: neg-reject 57.33 vs the peer's 39.0** — this model declines to answer unanswerable questions far more often than it invents an answer. Noise-robustness (64.0) improved over an earlier revision (58.67) but remains below the 70 threshold we treat as a pass. ### IFEval in context (strict-avg) | model | weights | company | params | IFEval strict-avg | | :-- | :-- | :-- | :-- | --: | | GPT-4o-mini | Proprietary | OpenAI | 8B + | 79 \* | | Llama-3.2-3B-Instruct | Open | Meta | 3.2B | 73.9 | | **KOS-V5-Instruct (ours)** | **Open** | **Univ. of Kentucky / Louisville** | **3.7B** | **72.19** | | Qwen2.5-3B-Instruct | Open | Alibaba | 3.0B | 64.7 | | Phi-3-medium-4k-instruct | Open | Microsoft | 14.0B | 64.2 | | Mistral-Large | Proprietary | Mistral AI | 46.7B + | 63 \* | | **KOS-V4-Instruct (previous gen)** | **Open** | **Univ. of Kentucky** | **3.0B** | **61.6** | | Yi-1.5-9B-Chat | Open | 01.AI | 8.8B | 60.5 | | Phi-3.5-mini-instruct | Open | Microsoft | 3.8B | 57.7 | | GPT-3.5-turbo-0613 | Proprietary | OpenAI | 20B + | 57 \* | | Phi-3-mini-4k-instruct | Open | Microsoft | 3.8B | 56.1 | | GPT-3.5-turbo-1106 | Proprietary | OpenAI | 20B + | 55.9 | | Mistral-7B-Instruct-v0.2 | Open | Mistral AI | 7.2B | 55.0 | | Llama-3.1-8B-Instruct | Open | Meta | 8.0B | 44.3 | | Llama-2-13b-chat | Open | Meta | 13.0B | 39.8 | **\* strict estimate** — no official IFEval strict sub-metrics published; estimated from published AVG4 or prompt-strict (loose metrics run ~2–4 pts above strict). **\+ unofficial params.** ### Against university-built instruction models Measured by us on the identical pristine harness, same protocol: | model | institution | params | IFEval strict-avg | MMLU | | :-- | :-- | --: | --: | --: | | **KOS-V5-Instruct (ours)** | **UK / UofL** | **3.7B** | **72.19** | **0.4512** | | marin-8b-instruct | Stanford | 8.0B | 70.83 | 0.6112 | | YuLan-Mini-Instruct | Renmin | 2.4B | 61.51 | 0.5278 | | **KOS-V4-Instruct (ours)** | **UK** | **3.0B** | **60.63** | 0.2782 | | LLäMmlein-7B-chat | Würzburg | 7.0B | 54.07 | 0.5252 | | Poro-34B-chat | U Turku | 34.2B | 34.63 | — | | Minerva-7B-instruct | Sapienza | 7.4B | 21.51 | 0.4071 | | CroissantLLMChat | CentraleSupélec | 1.3B | 19.94 | 0.2401 | | Tucano-2b4-Instruct | U Bonn | 2.4B | 14.95 | 0.2589 | On instruction following KOS-V5-Instruct now places **first among nine university-built instruct models**, ahead of Stanford's Marin-8B (70.83) at less than half its parameters, and of Poro-34B at **9× its parameter count**. Note that several of these models are non-English-first (Finnish, Italian, French, Portuguese, German) and are being measured on English benchmarks, which understates their designed capability. On parametric knowledge (MMLU) the larger, more heavily trained models still lead. ## Retrieval & embeddings Beyond generation, KOS-V5-Instruct **also serves as a dense text retriever.** A companion **LoRA adapter** — [**KOS-V5-Retriever**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Retriever) — converts this model into an embedding model (llm2vec-style: bidirectional attention + mean-pooling + a contrastively-trained rank-32 LoRA), with these base weights **frozen and unchanged**. On the **official BEIR SciFact** benchmark (the `beir` library + pytrec_eval — the public-leaderboard scorer), **zero-shot** (training excluded SciFact, verified clean), it scores **NDCG@10 = 0.7007** (Recall@10 0.864) — a **strong dense retriever**, above BM25 (0.665) and within the GTR/E5/BGE band (0.70–0.76). The adapter is hot-swappable: attach it for retrieval, detach it for generation. See the [adapter card](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Retriever) for the encode recipe and the full retrieval details. ## Known regressions and limitations > These are disclosed deliberately. A high benchmark score does not make this checkpoint production-ready. - **Tool / function calling is a trained strength — and now measured against the peer.** The second GRPO leg optimised the *official* BFCL AST checker directly; BFCL is 85.0/84.0/80.5, above the peer and far above the previous KOS generation. Earlier internal KOS tool-calling figures are deliberately omitted: several were measured against benchmark data the model had been trained on and are recorded in our own audit as **invalid**. - **Grounded abstention / fabrication IS measured — and strong.** On the official RGB harness, negative-rejection is 57.33 (peer 39.0). *Earlier revisions of this card stated fabrication was NOT independently measured; it now is.* Noise-robustness (64.0) is still below our 70 pass bar, so retrieval-noise handling remains a known gap. - **The peer leads on knowledge and raw instruction following.** IFEval 72.19 vs the peer's 84.71; MMLU 0.4512 vs 0.7266; medical QA below the peer. This is a from-scratch 3.7B model on 235B tokens against one trained on orders of magnitude more data — strong for its scale, not state-of-the-art in absolute terms. - **MMLU 0.4512 is above chance (0.25) but modest.** This is not a knowledge model; it should not be used as a medical question-answering authority. - **Not a medical-MCQ model.** As with KOS-V4, do not benchmark or deploy it as one; read the BPB results (base card) as the medical signal and the MCQ results as a floor. ## Data contamination - **IFEval: CLEAN (verbatim).** The SFT mix and the IFEval RL prompt set were exact-containment scanned against IFEval's official 541 test prompts — **0 exact containments**. Exact matching cannot detect paraphrase or reformatting. - **Tool calling (BFCL): CLEAN (verbatim).** The tool-calling RL prompt pool was exact-containment scanned against **5,437 full-length official BFCL prompts — 0 exact containments**. - **MMLU / PubMedQA / MedQA / MedMCQA: UNCHECKED.** Contamination against these four has **not** been scanned for this checkpoint. Those numbers should be read with that caveat. ## Prompt / chat format (ChatML) ``` <|im_start|>system You are a helpful assistant.<|im_end|> <|im_start|>user Write a haiku about Kentucky. Do not use any commas.<|im_end|> <|im_start|>assistant ``` Tool / function calling uses the model's native `` … `` schema block and `` … `` response format; pass your function schemas via the tokenizer's `apply_chat_template(..., tools=[...])`. ## Quickstart ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "Kentucky-Open-Science/KOS-V5-Instruct" tok = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto") messages = [{"role": "user", "content": "Write a haiku about Kentucky. Do not use any commas."}] inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device) out = model.generate(inputs, max_new_tokens=256, do_sample=False) print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True)) ``` ## Deployment | Precision | Approx. VRAM | Notes | | :--- | :--- | :--- | | **bfloat16** | ~9 GB | native weights (7.43 GB) + activations; a single 16 GB GPU is comfortable | ## Related models - [**KOS-V5-Retriever**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Retriever) — a LoRA **retrieval/embedding adapter** for this model (llm2vec; official BEIR SciFact NDCG@10 **0.70**, zero-shot). - [**KOS-V5-Base**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) — the from-scratch pretrained foundation this model is tuned from (3.72B, 235.2B tokens). - [**KOS-V4-Instruct**](https://huggingface.co/Kentucky-Open-Science/KOS-V4-Instruct) — previous generation (3.0B, IFEval 61.6), the public release. - [**KOS-V4-Base**](https://huggingface.co/Kentucky-Open-Science/KOS-V4-Base) — previous-generation foundation (3.015B, 180.3B tokens). ## Intended use & limitations Research use only. English only. Not for clinical, commercial, legal, or production-grade use. Outputs may be factually wrong or fabricated. This model must not be used to make or inform medical decisions. ## Naming The program is **KOS** (KOS-V1..V6). Earlier internal names are not used.