--- base_model: meta-llama/Llama-3.1-8B-Instruct library_name: peft license: llama3.1 tags: - lora - peft - knowledge-graph-completion - link-prediction - biomedical - cold-start pipeline_tag: text-generation --- # Cold-Start Chemical–Gene Ranking — LoRA adapter (8B) **Built with Llama.** This is a LoRA (r=64) adapter for **`meta-llama/Llama-3.1-8B-Instruct`**, from the paper *"Cold-Start Link Prediction Needs a Ranking Readout"* (GaLM 2026 @ CIKM). The fine-tuned model is read out as a **length-normalized sequence log-probability scorer** to rank candidate genes for a chemical–gene interaction query, under a cold-start (unseen-chemical) protocol. ## What this is - **Adapter only** (~640 MB). The base model `meta-llama/Llama-3.1-8B-Instruct` is **not** included — download it from the Hugging Face Hub (subject to the Llama Community License). - Fine-tuned on a **CTD-derived** chemical–gene QA corpus (augmented *sample-47* footing). - Part of a size ladder released with the paper — **1B 0.78 / 3B 0.86 / 8B 0.92** (cold-start, sampled hard-negative MRR, K=99, sample-47). Under this capacity-limited LoRA regime, the readout improves monotonically with size, which — together with the full-fine-tuning result where a 1B already reaches the ceiling — locates the operative axis at **capacity, not scale**. ## This adapter - **Cold-start sampled hard-negative MRR = 0.917** (ep8; sample-47 corpus, K=99). - LoRA config: r=64, alpha=64, target modules q/k/v/o/gate/up/down_proj. ## Intended use & limitations - **Research use only.** Cold-start chemical–gene *ranking* (scoring), not free generation and not clinical decision-making. Absolute values are only comparable **within** the sample-47, LoRA footing (never cross-compared with the full-fine-tuning / gl47 headline numbers). ## Usage (sketch) ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig from peft import PeftModel base_id = "meta-llama/Llama-3.1-8B-Instruct" tok = AutoTokenizer.from_pretrained("BioRel/coldstart-lora-8b") bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_quant_type="nf4") base = AutoModelForCausalLM.from_pretrained(base_id, quantization_config=bnb, device_map={"": 0}) lm = PeftModel.from_pretrained(base, "BioRel/coldstart-lora-8b").eval() # score a candidate gene g for query q by length-normalized log-prob of g given q # s(q, g) = (1/|g|) * sum_t log P(g_t | q, g_