Text Generation
PEFT
Safetensors
lora
knowledge-graph-completion
link-prediction
biomedical
cold-start
conversational
Instructions to use BioRel/coldstart-lora-8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use BioRel/coldstart-lora-8b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct") model = PeftModel.from_pretrained(base_model, "BioRel/coldstart-lora-8b") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from BioRel/coldstart-lora-8b: direct link, hf CLI and curl.
- Browser
- Download file 3.72 kB
-
https://huggingface.co/BioRel/coldstart-lora-8b/resolve/main/README.md
- Command line
-
hf download hf://BioRel/coldstart-lora-8b/README.md
-
curl -L -o README.md https://huggingface.co/BioRel/coldstart-lora-8b/resolve/main/README.md
3.72 kB
| base_model: meta-llama/Llama-3.1-8B-Instruct | |
| library_name: peft | |
| license: llama3.1 | |
| tags: | |
| - lora | |
| - peft | |
| - knowledge-graph-completion | |
| - link-prediction | |
| - biomedical | |
| - cold-start | |
| pipeline_tag: text-generation | |
| # Cold-Start Chemical–Gene Ranking — LoRA adapter (8B) | |
| **Built with Llama.** This is a LoRA (r=64) adapter for **`meta-llama/Llama-3.1-8B-Instruct`**, from the paper | |
| *"Cold-Start Link Prediction Needs a Ranking Readout"* (GaLM 2026 @ CIKM). The fine-tuned | |
| model is read out as a **length-normalized sequence log-probability scorer** to rank candidate | |
| genes for a chemical–gene interaction query, under a cold-start (unseen-chemical) protocol. | |
| ## What this is | |
| - **Adapter only** (~640 MB). The base model `meta-llama/Llama-3.1-8B-Instruct` is **not** included — download it from the | |
| Hugging Face Hub (subject to the Llama Community License). | |
| - Fine-tuned on a **CTD-derived** chemical–gene QA corpus (augmented *sample-47* footing). | |
| - Part of a size ladder released with the paper — **1B 0.78 / 3B 0.86 / 8B 0.92** (cold-start, | |
| sampled hard-negative MRR, K=99, sample-47). Under this capacity-limited LoRA regime, the | |
| readout improves monotonically with size, which — together with the full-fine-tuning result | |
| where a 1B already reaches the ceiling — locates the operative axis at **capacity, not scale**. | |
| ## This adapter | |
| - **Cold-start sampled hard-negative MRR = 0.917** (ep8; sample-47 corpus, K=99). | |
| - LoRA config: r=64, alpha=64, target modules q/k/v/o/gate/up/down_proj. | |
| ## Intended use & limitations | |
| - **Research use only.** Cold-start chemical–gene *ranking* (scoring), not free generation and | |
| not clinical decision-making. Absolute values are only comparable **within** the sample-47, | |
| LoRA footing (never cross-compared with the full-fine-tuning / gl47 headline numbers). | |
| ## Usage (sketch) | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig | |
| from peft import PeftModel | |
| base_id = "meta-llama/Llama-3.1-8B-Instruct" | |
| tok = AutoTokenizer.from_pretrained("BioRel/coldstart-lora-8b") | |
| bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_quant_type="nf4") | |
| base = AutoModelForCausalLM.from_pretrained(base_id, quantization_config=bnb, device_map={"": 0}) | |
| lm = PeftModel.from_pretrained(base, "BioRel/coldstart-lora-8b").eval() | |
| # score a candidate gene g for query q by length-normalized log-prob of g given q | |
| # s(q, g) = (1/|g|) * sum_t log P(g_t | q, g_<t) ; rank genes by s. | |
| # Full scorer: https://github.com/BioRel/relational-qa-coldstart (src/06_baselines/lora_scorer.py) | |
| ``` | |
| ## Training data & license | |
| - **Data:** derived from the **Comparative Toxicogenomics Database (CTD)**, https://ctdbase.org. | |
| The dataset is **subject to CTD terms**; users must download CTD data themselves | |
| (terms: https://ctdbase.org/about/legal.jsp). **CTD may access this dataset for quality control | |
| purposes.** Non-commercial / research use. | |
| - **Base model:** `meta-llama/Llama-3.1-8B-Instruct` — governed by the **Llama Community License** (llama3.1). You must accept | |
| Meta's license to download the base. | |
| - **Adapter weights:** released for research use, subject to the base-model and CTD terms above. | |
| ## Citation | |
| ```bibtex | |
| @inproceedings{kim2026coldstart, | |
| title = {Cold-Start Link Prediction Needs a Ranking Readout}, | |
| author = {Kim, Yunha and Kim, Young-Hak and Jun, Tae Joon}, | |
| booktitle = {Proceedings of the Workshop on Graph-Augmented LLMs (GaLM), co-located with CIKM}, | |
| year = {2026} | |
| } | |
| ``` | |
| Please also cite CTD: A. P. Davis et al., *Comparative Toxicogenomics Database (CTD): update 2021*, | |
| Nucleic Acids Research, 2021. https://ctdbase.org | |