Text Generation
Transformers
Safetensors
English
qwen3
medical
clinical
biomedical
instruction-following
tool-calling
function-calling
from-scratch
KOS-V5
university-of-kentucky
university-of-louisville
ccts
conversational
text-generation-inference
Instructions to use Kentucky-Open-Science/KOS-V5-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Kentucky-Open-Science/KOS-V5-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Kentucky-Open-Science/KOS-V5-Instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Kentucky-Open-Science/KOS-V5-Instruct") model = AutoModelForCausalLM.from_pretrained("Kentucky-Open-Science/KOS-V5-Instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Kentucky-Open-Science/KOS-V5-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Kentucky-Open-Science/KOS-V5-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kentucky-Open-Science/KOS-V5-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Kentucky-Open-Science/KOS-V5-Instruct
- SGLang
How to use Kentucky-Open-Science/KOS-V5-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Kentucky-Open-Science/KOS-V5-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kentucky-Open-Science/KOS-V5-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Kentucky-Open-Science/KOS-V5-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kentucky-Open-Science/KOS-V5-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Kentucky-Open-Science/KOS-V5-Instruct with Docker Model Runner:
docker model run hf.co/Kentucky-Open-Science/KOS-V5-Instruct
|
Download README.md from Kentucky-Open-Science/KOS-V5-Instruct: direct link, hf CLI and curl.
- Browser
- Download file 22.2 kB
-
https://huggingface.co/Kentucky-Open-Science/KOS-V5-Instruct/resolve/main/README.md
- Command line
-
hf download hf://Kentucky-Open-Science/KOS-V5-Instruct/README.md
-
curl -L -o README.md https://huggingface.co/Kentucky-Open-Science/KOS-V5-Instruct/resolve/main/README.md
22.2 kB
| license: apache-2.0 | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| developed_by: University of Kentucky (College of Medicine Office for Research; Center for Clinical and Translational Science) and University of Louisville (Kentucky Center for Digital Innovation) | |
| affiliations: | |
| - name: University of Kentucky, College of Medicine Office for Research | |
| url: https://medicine.uky.edu/sites/research | |
| - name: University of Kentucky, Center for Clinical and Translational Science (CCTS) | |
| url: https://www.ccts.uky.edu/ | |
| - name: University of Louisville, Kentucky Center for Digital Innovation | |
| url: https://centers.louisville.edu/kentucky-center-digital-innovation | |
| model_name: KOS-V5-Instruct | |
| model_codename: Catbird | |
| model_type: qwen3 | |
| base_model: Kentucky-Open-Science/KOS-V5-Base | |
| tags: | |
| - medical | |
| - clinical | |
| - biomedical | |
| - instruction-following | |
| - tool-calling | |
| - function-calling | |
| - from-scratch | |
| - qwen3 | |
| - KOS-V5 | |
| - university-of-kentucky | |
| - university-of-louisville | |
| - ccts | |
| <p align="center"> | |
| <img src="catbird_llm_logo.png" alt="Catbird" width="320"/> | |
| </p> | |
| # KOS-V5-Instruct · *"Catbird"* | |
| **Developed by** | |
| **University of Kentucky** | |
| - [College of Medicine, Office for Research](https://medicine.uky.edu/sites/research) | |
| - [Center for Clinical and Translational Science (CCTS)](https://www.ccts.uky.edu/) | |
| **University of Louisville** | |
| - [Kentucky Center for Digital Innovation](https://centers.louisville.edu/kentucky-center-digital-innovation) | |
| **A 3.72B-parameter medical language model trained from scratch.** It is not distilled, not pruned and not | |
| continued-pretrained from a general base. KOS-V5 (codename **Catbird**) is the fifth-generation Kentucky Open | |
| Science model line. This repository holds the **instruction-tuned head** of that line: the | |
| [KOS-V5-Base](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) pretraining checkpoint, taken through | |
| SFT and two GRPO reinforcement-learning legs. | |
| Unlike the base, this model **follows instructions and calls tools**. It is the downstream SFT/RL artifact that | |
| [KOS-V5-Base](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) was built to initialise. | |
| **Code name: Catbird.** The KOS-V5 series is nicknamed Catbird; native to Kentucky, the Gray Catbird | |
| (*Dumetella carolinensis*) is a medium-sized songbird famous for its distinct, cat-like "meow" call. This LLM | |
| was trained completely from scratch by teams from the University of Kentucky (Cat) and University of Louisville | |
| (Bird), so the code name is fitting. | |
| > ⚠️ **Research use only.** This model is provided for research purposes only and must not be used for any | |
| > commercial, clinical, legal, or production-grade application. The user assumes all risks associated with its use. | |
| --- | |
| Its instruction ability comes from **GRPO reinforcement learning** against the *official* IFEval verifier, and its | |
| tool-calling ability from a **second GRPO leg** against the *official* BFCL AST checker, on a base that ranks | |
| **first of 17** at modelling held-out clinical text. | |
| **IFEval reported as strict-avg** = `(prompt-level strict + instruction-level strict) / 2` — the exact metric the | |
| Hugging Face Open LLM Leaderboard publishes as "IFEval." | |
| | IFEval **strict-avg** | model | who built it, and how | | |
| | --: | :-- | :-- | | |
| | **72.19** | **KOS-V5-Instruct (ours)** | University research team, 235B tokens, from scratch | | |
| | 64.7 | Qwen2.5-3B-Instruct | Alibaba, ~18 trillion tokens | | |
| | **61.6** | [KOS-V4-Instruct](https://huggingface.co/Kentucky-Open-Science/KOS-V4-Instruct) (previous generation) | University research team, 180B tokens, 24 GPUs | | |
| | 55.9 | GPT-3.5-turbo-1106 (the original ChatGPT) | OpenAI, ~10,000-GPU supercomputer | | |
| KOS-V5-Instruct **improves on KOS-V4-Instruct across every benchmark measured**: IFEval strict-avg | |
| **61.6 → 72.19** (+10.6), MMLU **0.2782 → 0.4512** (+17.3), medical QA (**PubMedQA 0.7060**, MedQA 0.3802, | |
| MedMCQA 0.3648 — all up on V4), and official BFCL function-calling | |
| **72.75/73.00/60.50 → 85.00/84.00/80.50** (+12.3 / +11.0 / +20.0). It clears the original GPT-3.5-turbo | |
| generation and the commercially trained Qwen2.5-3B on instruction following, and its tool calling now runs | |
| **above the Qwen3-4B-Instruct-2507 peer**. | |
| ## Core specifications | |
| | Attribute | Detail | | |
| | :--- | :--- | | |
| | **Architecture** | Decoder-only Transformer (`Qwen3ForCausalLM`), Grouped-Query Attention | | |
| | **Parameters** | 3.715 B | | |
| | **Hidden / Layers** | 2560 / 36 | | |
| | **Attention** | 32 query / 8 KV heads (GQA 4:1), head_dim 128, per-head QK-RMSNorm | | |
| | **Feed-forward** | SwiGLU, intermediate 9728 | | |
| | **Vocabulary** | 32,000, custom medical byte-level BPE | | |
| | **Context length** | 32,768 | | |
| | **Position encoding** | RoPE, θ = 25,000 | | |
| | **Embeddings** | tied | | |
| | **Precision** | bfloat16 (7.43 GB, single shard) | | |
| ## Pre-training (the KOS-V5 base) | |
| Fine-tuned from [**KOS-V5-Base**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) — the from-scratch | |
| pretrained foundation, a complete single-epoch run over **235.2B tokens**. See that card for corpus composition and | |
| disclosed pretraining issues. | |
| ## Post-training (this model) | |
| Three stages on top of the base — no LoRA, no distillation, no reward model, no LLM judge. | |
| **SFT** — one shuffled full-parameter pass over a **736,990-record / 1.32B-token** audited instruction mix | |
| (ChatML). The mix was deduplicated, instruction-collision canonicalized, structurally validated, per-record | |
| decontaminated and BFCL-closure scanned; clinical FHIR records were dropped and tool-record system prompts | |
| de-welded so tool use binds to the request rather than to a fixed frame. | |
| **RL leg 1 — instruction following (GRPO via verl)** — deterministic verifiable reward. The reward is the | |
| **official** `lm_eval` IFEval instruction registry — the same checkers the benchmark scores with, not a | |
| re-implementation. One 8-GPU wall, 164 steps, KL 0.001 (`low_var_kl`), rollout n=16. | |
| **RL leg 2 — tool calling (GRPO via verl)** — a second leg seeded from leg 1. The reward is the **official** | |
| BFCL `ast_checker` (`bfcl_eval`). Each prompt renders its tool schemas through the model's **own** chat template | |
| (byte-exact to the official `tools=` rendering), and the prompt set is filtered to only rows the official checker | |
| can grade. 8-GPU wall, KL 0.001, rollout n=16; this repository ships the **step-40** checkpoint, selected for the | |
| best BFCL / abstention balance and least policy drift. BFCL rose **76.8/71.0/69.5 → 85.0/84.0/80.5** with | |
| instruction following, grounded abstention and knowledge all held. | |
| **Forgetting control** — out-of-distribution broad-holdout perplexity at **0.99× the pre-RL base** (8.88 vs 8.97), | |
| measured on a web crawl postdating the training corpus. No measurable forgetting. | |
| ## The medical foundation | |
| **This is a medical model.** KOS-V5-Instruct inherits a base trained on a **54-source medical/biomedical corpus** | |
| — not a general-purpose model with medical fine-tuning bolted on. | |
| The strongest evidence is **bits-per-byte on held-out medical text**, which is tokenizer-agnostic and therefore | |
| the only strictly fair cross-model comparison. In a **17-model pool** — including dedicated biomedical | |
| specialists BioMedLM (300B PubMed tokens), Meditron-7B, PMC-LLaMA-7B and MedGemma-4B — the KOS-V5 base ranks | |
| **first**: | |
| | medical text (BPB, lower is better) | KOS-V5-Base | rank | | |
| | :-- | --: | --: | | |
| | **5-corpus mean, held-out medical text** | **0.4635** | **1 / 17** | | |
| | clinical narratives | **0.4179** | **1 / 17** | | |
| | radiology | **0.5132** | **1 / 17** | | |
| | chest X-ray reports | **0.6688** | **1 / 17** | | |
| | BIOSSES biomedical sentence similarity (Pearson / Spearman) | **0.7097 / 0.7014** | **1 / 17** | | |
| | BLURB biomedical probe mean | 0.7268 | 2 / 17 | | |
| Every comparator in that pool was trained on **1.3–153× more data** (0.3–36T tokens vs our 0.235T). See | |
| [KOS-V5-Base](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) for the full 96-metric evaluation. | |
| ### Medical MMLU (from the 57-subject run above) | |
| The 9 medical subjects of MMLU, extracted from the same official 5-shot run: | |
| | medical subject | KOS-V5-Instruct | KOS-V4-Instruct | | |
| | :-- | --: | --: | | |
| | high-school biology | 0.5774 | 0.2387 | | |
| | clinical knowledge | 0.5623 | 0.3170 | | |
| | nutrition | 0.5359 | 0.2843 | | |
| | college biology | 0.5347 | 0.2917 | | |
| | medical genetics | 0.5100 | 0.2700 | | |
| | anatomy | 0.4815 | 0.3185 | | |
| | professional medicine | 0.4375 | 0.2132 | | |
| | college medicine | 0.4046 | 0.2486 | | |
| | virology | 0.3795 | 0.2952 | | |
| | **medical-9 mean** | **0.4915** | **0.2752** | | |
| **+21.6 points over KOS-V4-Instruct**, and above the model's own full-MMLU average (0.4512) — the medical | |
| subjects are where it is strongest. | |
| ### Medical QA benchmarks (official suites) | |
| Official `medqa_4options`, `medmcqa` and `pubmedqa` from the same pristine lm-evaluation-harness, | |
| 5-shot, loglikelihood, metric `acc`. | |
| | medical benchmark | **KOS-V5-Instruct** | KOS-V4-Instruct | YuLan-Mini-Instruct | marin-8b-instruct | Qwen3-4B-Instruct-2507 | | |
| | :-- | --: | --: | --: | --: | --: | | |
| | params | **3.7B** | 3.0B | 2.4B | 8.0B | 4.0B | | |
| | **PubMedQA** | **0.7060** | 0.6860 | 0.6960 | 0.7500 | 0.7720 | | |
| | MedQA (USMLE, 4-option) | **0.3802** | 0.2820 | 0.3511 | 0.4878 | 0.6159 | | |
| | MedMCQA | **0.3648** | 0.2778 | 0.3856 | 0.4961 | 0.5804 | | |
| **KOS-V5-Instruct improves on KOS-V4-Instruct on all three** (+9.8 MedQA, +8.7 MedMCQA, +2.0 PubMedQA). | |
| **PubMedQA is the standout: 0.7060**, ahead of YuLan-Mini and within reach of Stanford's Marin-8B at **less than | |
| half the parameters**. PubMedQA tests comprehension of **biomedical literature** — the closest of these three to | |
| what the base was actually trained on. The USMLE-style exam MCQs (MedQA, MedMCQA) are where the answer-letter | |
| bottleneck below bites hardest. | |
| > ⚠️ **Why the MCQ numbers understate this model.** Our own measurements show KOS models place very little | |
| > probability mass on MCQ answer *letters*: the format, not the knowledge, is the bottleneck. A model that ranks | |
| > **1 of 17** at modelling clinical text while scoring modestly on multiple-choice is exhibiting exactly that gap. | |
| > **Read the BPB results as the medical signal and the MCQ results as a floor, not a ceiling.** | |
| ## Evaluation & benchmarks | |
| **Official suites only**, EleutherAI lm-evaluation-harness `0.4.12.dev0` at upstream commit `c1c4bea`, run from a | |
| **pristine clone** with stock, unmodified task definitions. | |
| - **IFEval** — stock `ifeval` task, 0-shot, greedy (`do_sample=false`, `temperature=0.0`), | |
| task-default `max_gen_toks=1280`, `--apply_chat_template`, seed 0. Constraint checking by the harness's vendored | |
| **Google** verifier (`instructions_registry`, 25 instruction types). | |
| - **MMLU** — stock `mmlu` group, official 57 subjects / 14,042 test items, 5-shot from `dev` (`first_n`), | |
| loglikelihood over A–D, metric `acc` (not `acc_norm`), **no chat template**. | |
| - **Medical QA** — stock `medqa_4options`, `medmcqa`, `pubmedqa` tasks, 5-shot, loglikelihood, metric `acc`, | |
| no chat template. | |
| | benchmark | KOS-V5-Instruct | KOS-V4-Instruct | Δ | | |
| | :-- | --: | --: | --: | | |
| | **IFEval strict-avg** | **72.19** | 61.6 | **+10.6** | | |
| | IFEval prompt-strict | 0.6728 | 0.5471 | +0.126 | | |
| | IFEval inst-strict | 0.7710 | 0.6655 | +0.106 | | |
| | IFEval prompt-loose | 0.6932 | 0.5693 | +0.124 | | |
| | IFEval inst-loose | 0.7878 | 0.6882 | +0.100 | | |
| | **MMLU (57-subj, 5-shot, `acc`)** | **0.4512** | 0.2782 | **+0.173** | | |
| ### Tool / function calling — official BFCL | |
| Measured with the **official `bfcl_eval`** suite in **FC (function-calling) mode**, non-live categories, the model | |
| prompted in its own native tool format and served via vLLM. **Tool calling is a trained objective of this model** | |
| — the second GRPO leg optimised the official BFCL AST checker directly. | |
| | BFCL (official, FC mode, **non-live AST**) | **KOS-V5-Instruct** | Qwen3-4B-Instruct-2507 (peer) | KOS-V4-Instruct | | |
| | :-- | --: | --: | --: | | |
| | simple *(334/400)* | **85.00** | 83.20 | 72.75 | | |
| | multiple *(157/200)* | **84.00** | 79.00 | 73.00 | | |
| | parallel *(147/200)* | **80.50** | 73.50 | 60.50 | | |
| **KOS-V5-Instruct is above the Qwen3-4B-Instruct-2507 peer on all three BFCL categories**, and far above the | |
| previous KOS-V4-Instruct. This is the axis the tool-calling GRPO leg was built to move, and it moved. | |
| > **Scope.** These are the **non-live AST** categories only (`simple_python`, `multiple`, `parallel`). The | |
| > live, multi-turn, web-search and memory categories were **not run**, so no BFCL *overall* score is reported | |
| > here — the suite's aggregate column is not meaningful when most categories are unrun. | |
| > **Engine note.** These BFCL numbers come from the official `bfcl_eval` harness on a **vLLM** backend, whereas | |
| > the IFEval and MMLU figures on this card come from the HuggingFace backend of a pristine lm-evaluation-harness. | |
| > Both are official suites; they are not the same inference stack, and that is stated rather than blurred. | |
| **Cross-harness reproduction.** IFEval strict-avg measured **72.19** (pristine HF harness) and **72.0** (our | |
| RL-evaluation harness) in two independent runs — a **0.19-point** agreement across two harness builds, far below | |
| the benchmark's own ±2.14-point standard error on 541 prompts, so they are the same measurement. | |
| **Harness validation.** The identical pipeline scored the peer mark **Qwen3-4B-Instruct-2507 at 84.71** IFEval | |
| strict-avg on the same pristine harness, and independently reproduced KOS-V4-Instruct's MMLU to four decimal | |
| places (0.2782). A score of 0.0 on this pipeline would therefore be a model property, not a harness failure. | |
| ### Grounded abstention & robustness — official RGB | |
| Measured on the **official RGB harness** (retrieval-augmented generation benchmark). | |
| | RGB (official) | KOS-V5-Instruct | Qwen3-4B-Instruct-2507 (peer) | | |
| | :-- | --: | --: | | |
| | **negative rejection** (declines the unanswerable) | **57.33** | 39.0 | | |
| | noise robustness | 64.0 | 93.67 | | |
| **Grounded abstention is a genuine strength: neg-reject 57.33 vs the peer's 39.0** — this model declines to answer | |
| unanswerable questions far more often than it invents an answer. Noise-robustness (64.0) improved over an earlier | |
| revision (58.67) but remains below the 70 threshold we treat as a pass. | |
| ### IFEval in context (strict-avg) | |
| | model | weights | company | params | IFEval strict-avg | | |
| | :-- | :-- | :-- | :-- | --: | | |
| | GPT-4o-mini | Proprietary | OpenAI | 8B + | 79 \* | | |
| | Llama-3.2-3B-Instruct | Open | Meta | 3.2B | 73.9 | | |
| | **KOS-V5-Instruct (ours)** | **Open** | **Univ. of Kentucky / Louisville** | **3.7B** | **72.19** | | |
| | Qwen2.5-3B-Instruct | Open | Alibaba | 3.0B | 64.7 | | |
| | Phi-3-medium-4k-instruct | Open | Microsoft | 14.0B | 64.2 | | |
| | Mistral-Large | Proprietary | Mistral AI | 46.7B + | 63 \* | | |
| | **KOS-V4-Instruct (previous gen)** | **Open** | **Univ. of Kentucky** | **3.0B** | **61.6** | | |
| | Yi-1.5-9B-Chat | Open | 01.AI | 8.8B | 60.5 | | |
| | Phi-3.5-mini-instruct | Open | Microsoft | 3.8B | 57.7 | | |
| | GPT-3.5-turbo-0613 | Proprietary | OpenAI | 20B + | 57 \* | | |
| | Phi-3-mini-4k-instruct | Open | Microsoft | 3.8B | 56.1 | | |
| | GPT-3.5-turbo-1106 | Proprietary | OpenAI | 20B + | 55.9 | | |
| | Mistral-7B-Instruct-v0.2 | Open | Mistral AI | 7.2B | 55.0 | | |
| | Llama-3.1-8B-Instruct | Open | Meta | 8.0B | 44.3 | | |
| | Llama-2-13b-chat | Open | Meta | 13.0B | 39.8 | | |
| **\* strict estimate** — no official IFEval strict sub-metrics published; estimated from published AVG4 or | |
| prompt-strict (loose metrics run ~2–4 pts above strict). **\+ unofficial params.** | |
| ### Against university-built instruction models | |
| Measured by us on the identical pristine harness, same protocol: | |
| | model | institution | params | IFEval strict-avg | MMLU | | |
| | :-- | :-- | --: | --: | --: | | |
| | **KOS-V5-Instruct (ours)** | **UK / UofL** | **3.7B** | **72.19** | **0.4512** | | |
| | marin-8b-instruct | Stanford | 8.0B | 70.83 | 0.6112 | | |
| | YuLan-Mini-Instruct | Renmin | 2.4B | 61.51 | 0.5278 | | |
| | **KOS-V4-Instruct (ours)** | **UK** | **3.0B** | **60.63** | 0.2782 | | |
| | LLäMmlein-7B-chat | Würzburg | 7.0B | 54.07 | 0.5252 | | |
| | Poro-34B-chat | U Turku | 34.2B | 34.63 | — | | |
| | Minerva-7B-instruct | Sapienza | 7.4B | 21.51 | 0.4071 | | |
| | CroissantLLMChat | CentraleSupélec | 1.3B | 19.94 | 0.2401 | | |
| | Tucano-2b4-Instruct | U Bonn | 2.4B | 14.95 | 0.2589 | | |
| On instruction following KOS-V5-Instruct now places **first among nine university-built instruct models**, ahead of | |
| Stanford's Marin-8B (70.83) at less than half its parameters, and of Poro-34B at **9× its parameter count**. Note | |
| that several of these models are non-English-first (Finnish, Italian, French, Portuguese, German) and are being | |
| measured on English benchmarks, which understates their designed capability. On parametric knowledge (MMLU) the | |
| larger, more heavily trained models still lead. | |
| ## Retrieval & embeddings | |
| Beyond generation, KOS-V5-Instruct **also serves as a dense text retriever.** A companion **LoRA adapter** — | |
| [**KOS-V5-Retriever**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Retriever) — converts | |
| this model into an embedding model (llm2vec-style: bidirectional attention + mean-pooling + a contrastively-trained | |
| rank-32 LoRA), with these base weights **frozen and unchanged**. | |
| On the **official BEIR SciFact** benchmark (the `beir` library + pytrec_eval — the public-leaderboard scorer), | |
| **zero-shot** (training excluded SciFact, verified clean), it scores **NDCG@10 = 0.7007** (Recall@10 0.864) — a | |
| **strong dense retriever**, above BM25 (0.665) and within the GTR/E5/BGE band (0.70–0.76). The adapter is | |
| hot-swappable: attach it for retrieval, detach it for generation. See the | |
| [adapter card](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Retriever) for the encode recipe and | |
| the full retrieval details. | |
| ## Known regressions and limitations | |
| > These are disclosed deliberately. A high benchmark score does not make this checkpoint production-ready. | |
| - **Tool / function calling is a trained strength — and now measured against the peer.** The second GRPO leg | |
| optimised the *official* BFCL AST checker directly; BFCL is 85.0/84.0/80.5, above the peer and far above the | |
| previous KOS generation. Earlier internal KOS tool-calling figures are deliberately omitted: several were | |
| measured against benchmark data the model had been trained on and are recorded in our own audit as **invalid**. | |
| - **Grounded abstention / fabrication IS measured — and strong.** On the official RGB harness, negative-rejection | |
| is 57.33 (peer 39.0). *Earlier revisions of this card stated fabrication was NOT independently measured; it now | |
| is.* Noise-robustness (64.0) is still below our 70 pass bar, so retrieval-noise handling remains a known gap. | |
| - **The peer leads on knowledge and raw instruction following.** IFEval 72.19 vs the peer's 84.71; MMLU 0.4512 vs | |
| 0.7266; medical QA below the peer. This is a from-scratch 3.7B model on 235B tokens against one trained on orders | |
| of magnitude more data — strong for its scale, not state-of-the-art in absolute terms. | |
| - **MMLU 0.4512 is above chance (0.25) but modest.** This is not a knowledge model; it should not be used as a | |
| medical question-answering authority. | |
| - **Not a medical-MCQ model.** As with KOS-V4, do not benchmark or deploy it as one; read the BPB results (base | |
| card) as the medical signal and the MCQ results as a floor. | |
| ## Data contamination | |
| - **IFEval: CLEAN (verbatim).** The SFT mix and the IFEval RL prompt set were exact-containment scanned against | |
| IFEval's official 541 test prompts — **0 exact containments**. Exact matching cannot detect paraphrase or | |
| reformatting. | |
| - **Tool calling (BFCL): CLEAN (verbatim).** The tool-calling RL prompt pool was exact-containment scanned against | |
| **5,437 full-length official BFCL prompts — 0 exact containments**. | |
| - **MMLU / PubMedQA / MedQA / MedMCQA: UNCHECKED.** Contamination against these four has **not** been scanned for | |
| this checkpoint. Those numbers should be read with that caveat. | |
| ## Prompt / chat format (ChatML) | |
| ``` | |
| <|im_start|>system | |
| You are a helpful assistant.<|im_end|> | |
| <|im_start|>user | |
| Write a haiku about Kentucky. Do not use any commas.<|im_end|> | |
| <|im_start|>assistant | |
| ``` | |
| Tool / function calling uses the model's native `<tools>` … `</tools>` schema block and `<tool_call>` … `</tool_call>` | |
| response format; pass your function schemas via the tokenizer's `apply_chat_template(..., tools=[...])`. | |
| ## Quickstart | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "Kentucky-Open-Science/KOS-V5-Instruct" | |
| tok = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto") | |
| messages = [{"role": "user", "content": "Write a haiku about Kentucky. Do not use any commas."}] | |
| inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device) | |
| out = model.generate(inputs, max_new_tokens=256, do_sample=False) | |
| print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True)) | |
| ``` | |
| ## Deployment | |
| | Precision | Approx. VRAM | Notes | | |
| | :--- | :--- | :--- | | |
| | **bfloat16** | ~9 GB | native weights (7.43 GB) + activations; a single 16 GB GPU is comfortable | | |
| ## Related models | |
| - [**KOS-V5-Retriever**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Retriever) — a LoRA | |
| **retrieval/embedding adapter** for this model (llm2vec; official BEIR SciFact NDCG@10 **0.70**, zero-shot). | |
| - [**KOS-V5-Base**](https://huggingface.co/Kentucky-Open-Science/KOS-V5-Base) — the from-scratch pretrained | |
| foundation this model is tuned from (3.72B, 235.2B tokens). | |
| - [**KOS-V4-Instruct**](https://huggingface.co/Kentucky-Open-Science/KOS-V4-Instruct) — previous generation | |
| (3.0B, IFEval 61.6), the public release. | |
| - [**KOS-V4-Base**](https://huggingface.co/Kentucky-Open-Science/KOS-V4-Base) — previous-generation foundation | |
| (3.015B, 180.3B tokens). | |
| ## Intended use & limitations | |
| Research use only. English only. Not for clinical, commercial, legal, or production-grade use. Outputs may be | |
| factually wrong or fabricated. This model must not be used to make or inform medical decisions. | |
| ## Naming | |
| The program is **KOS** (KOS-V1..V6). Earlier internal names are not used. | |