PEFT
Safetensors
English
code
python
lora
qwen2
code-generation
SathishKumar89 commited on
Commit
b6df816
·
verified ·
1 Parent(s): acc11b8

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +46 -0
README.md CHANGED
@@ -36,6 +36,52 @@ This model was fine-tuned as a learning project to demonstrate the full workflow
36
  | **Hardware** | Google Colab (NVIDIA T4, 16 GB VRAM) |
37
  | **Training time** | ~33 minutes |
38
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
39
  ## Prompt Format
40
 
41
  This model was trained with the following instruction format. Using the same format at inference time will give the best results:
 
36
  | **Hardware** | Google Colab (NVIDIA T4, 16 GB VRAM) |
37
  | **Training time** | ~33 minutes |
38
 
39
+ ## What Is This — A Model or an Adapter?
40
+
41
+ This repository contains a **LoRA adapter**, not a standalone model. Understanding the difference matters for how you load and use it.
42
+
43
+ ### The Two Artifacts
44
+
45
+ | | **Base Model** | **LoRA Adapter (this repo)** |
46
+ |---|---|---|
47
+ | **What it is** | The full pretrained neural network | A small set of trained weights that modify the base |
48
+ | **Size** | ~3 GB | ~74 MB |
49
+ | **Who made it** | The Qwen team | Me (SathishKumar89) |
50
+ | **Repo** | `Qwen/Qwen2.5-Coder-1.5B-Instruct` | `SathishKumar89/my-python-coder` |
51
+ | **Contains** | All model weights, tokenizer, config | Only adapter weights + config + tokenizer copy |
52
+ | **Loadable alone?** | ✅ Yes | ❌ No — needs the base model |
53
+
54
+ ### Why This Design?
55
+
56
+ Instead of retraining all ~1.5 billion parameters of the base model, **LoRA (Low-Rank Adaptation)** freezes the base model and only trains a tiny number of new parameters. This gives several advantages:
57
+
58
+ - **Tiny file size** — 74 MB vs. ~3 GB (a ~40× reduction)
59
+ - **Fast training** — minutes to hours instead of days
60
+ - **Runs on modest hardware** — a free Google Colab T4 GPU is enough
61
+ - **Easy to swap** — you can keep the same base model and load different adapters for different tasks
62
+
63
+ ### How to Load It Correctly
64
+
65
+ Because this repo is an adapter, you must load **two** things — the base model first, then the adapter on top:
66
+
67
+ ```python
68
+ from transformers import AutoModelForCausalLM, AutoTokenizer
69
+ from peft import PeftModel
70
+ import torch
71
+
72
+ # Step 1: Load the base model
73
+ base = AutoModelForCausalLM.from_pretrained(
74
+ "Qwen/Qwen2.5-Coder-1.5B-Instruct",
75
+ dtype=torch.float16,
76
+ device_map="auto",
77
+ )
78
+
79
+ # Step 2: Attach the LoRA adapter
80
+ model = PeftModel.from_pretrained(base, "SathishKumar89/my-python-coder")
81
+
82
+ # Step 3: Load the tokenizer (included in this repo)
83
+ tokenizer = AutoTokenizer.from_pretrained("SathishKumar89/my-python-coder")
84
+
85
  ## Prompt Format
86
 
87
  This model was trained with the following instruction format. Using the same format at inference time will give the best results: