File size: 2,069 Bytes
cf47703
67915ce
9c1b53f
67915ce
 
 
 
 
9c1b53f
67915ce
 
9c1b53f
67915ce
 
 
 
cf47703
 
67915ce
cf47703
9c1b53f
 
 
cf47703
 
 
9c1b53f
67915ce
 
9c1b53f
 
67915ce
 
9c1b53f
cf47703
67915ce
cf47703
9c1b53f
cf47703
67915ce
 
 
cf47703
9c1b53f
cf47703
9c1b53f
 
 
 
67915ce
 
cf47703
67915ce
 
 
cf47703
67915ce
 
cf47703
9c1b53f
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
---
base_model: Qwen/Qwen2.5-0.5B-Instruct
library_name: transformers
pipeline_tag: text-generation
tags:
- code
- coding
- qwen
- qwen2
- slm
- trl
- fine-tuned
license: apache-2.0
language:
- en
- code
---

# light-coder

`light-coder` is an ultra-lightweight, standalone instruction-tuned coding model created by fine-tuning [Qwen/Qwen2.5-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct) on ~122k programming instruction-response pairs and merging the LoRA weights directly into the base checkpoint.

At under 1 GB in size, it requires minimal VRAM, executes quickly on consumer GPUs and CPUs, and integrates out-of-the-box with tools like vLLM, Ollama, and standard Hugging Face pipelines.

## Model Details

- **Developed by:** Milad Asghari
- **Model Name:** light-coder
- **Base Model:** [Qwen/Qwen2.5-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct)
- **Model Type:** Causal Language Model (Full Merged Weights)
- **Primary Domain:** Code generation, refactoring, and programming instruction-following
- **Language(s):** English, Multiple Programming Languages
- **License:** Apache-2.0
- **Size:** 988 MB (`safetensors`)

## How to Get Started

Because the weights are merged, you do not need the `peft` library for inference:

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Miladasghari/light-coder"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    dtype=torch.float16,
    device_map="auto"
)

messages = [
    {"role": "user", "content": "Write a Python function to check if a string is a palindrome."}
]

prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=256,
    temperature=0.3,
    top_p=0.9,
    repetition_penalty=1.05
)

response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response)