Cold-Start Chemical–Gene Ranking Readout (1B)
Built with Llama. Full fine-tune of meta-llama/Llama-3.2-1B-Instruct from the paper
"Cold-Start Link Prediction Needs a Ranking Readout" (GaLM 2026 @ CIKM). The model is read
out as a length-normalized sequence log-probability scorer to rank candidate genes for a
chemical–gene interaction query, under a cold-start (unseen-chemical) protocol.
This is the reproducible headline model: a generic (non-biomedical) Llama-3.2-1B, fine-tuned with a retained recipe, that reaches the same numbers as the biomedical 1B used in the paper — showing the result rides general pretraining, not domain-specific leakage.
Headline results (cold-start, full candidate set of 13,972 genes)
| Metric | Value |
|---|---|
| Full-candidate MRR | 0.536 |
| Sampled MRR (random neg, K=99) | 0.837 |
| Sampled MRR (hard neg, K=99) | 0.798 |
| Popularity floor (full-candidate) | 0.106 |
The score is s(h,r,g) = (1/|g|) · log P_θ(g | π(h,r)); genes are ranked by s.
Candidate-constrained decoding is equivalent to this ranking readout; unconstrained
generation fails at cold-start, which is the paper's central point.
Usage (sketch)
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "BioRel/coldstart-ranking-readout-1b"
tok = AutoTokenizer.from_pretrained(repo)
lm = AutoModelForCausalLM.from_pretrained(repo, torch_dtype=torch.bfloat16, device_map={"": 0}).eval()
# rank genes by length-normalized log-prob of each candidate given the (chemical, relation) question:
# s(q, g) = (1/|g|) * sum_t log P(g_t | q, g_<t)
# Full scorer: https://github.com/<repo> (src/06_baselines/path_b_llm_ranking.py, kg_llm_exactrank.py)
Intended use & limitations
- Research use only. Cold-start chemical–gene ranking (scoring), not free generation and not clinical decision-making.
- The companion LoRA adapters
BioRel/coldstart-lora-{1b,3b,8b}is also provided.
Training data & license
- Data: derived from the Comparative Toxicogenomics Database (CTD), https://ctdbase.org. Subject to CTD terms; users must download CTD data themselves (terms: https://ctdbase.org/about/legal.jsp). CTD may access this dataset for quality control purposes. Non-commercial / research use.
- Base model:
meta-llama/Llama-3.2-1B-Instruct— governed by the Llama 3.2 Community License. You must accept Meta's license. This fine-tune is distributed under the same terms.
Please cite CTD: A. P. Davis et al., Comparative Toxicogenomics Database (CTD): update 2021, Nucleic Acids Research, 2021. https://ctdbase.org
- Downloads last month
- 408
Model tree for BioRel/coldstart-ranking-readout-1b
Base model
meta-llama/Llama-3.2-1B-Instruct