File size: 9,245 Bytes
883e09c
f7b9ae7
 
 
 
 
 
 
 
 
 
 
 
 
883e09c
 
f7b9ae7
883e09c
f7b9ae7
 
 
883e09c
 
 
 
 
f7b9ae7
 
 
 
883e09c
f7b9ae7
 
 
 
 
883e09c
f7b9ae7
883e09c
f7b9ae7
883e09c
 
 
 
 
f7b9ae7
883e09c
f7b9ae7
 
 
 
 
 
 
 
 
 
 
 
883e09c
f7b9ae7
 
883e09c
 
 
f7b9ae7
 
 
 
 
 
883e09c
 
 
f7b9ae7
 
 
 
 
 
 
 
 
 
 
 
883e09c
 
 
f7b9ae7
 
 
883e09c
 
 
f7b9ae7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
883e09c
 
 
 
 
f7b9ae7
 
 
 
 
 
883e09c
 
 
f7b9ae7
 
 
 
 
883e09c
f7b9ae7
883e09c
f7b9ae7
 
 
883e09c
 
 
f7b9ae7
 
 
 
 
 
 
 
 
883e09c
f7b9ae7
883e09c
f7b9ae7
 
883e09c
 
 
 
 
 
 
f7b9ae7
 
883e09c
 
 
f7b9ae7
 
 
883e09c
 
 
f7b9ae7
 
 
 
 
 
883e09c
f7b9ae7
 
883e09c
f7b9ae7
883e09c
f7b9ae7
 
 
 
 
883e09c
 
 
f7b9ae7
 
 
 
 
 
883e09c
f7b9ae7
883e09c
 
 
f7b9ae7
 
 
 
883e09c
 
 
 
 
f7b9ae7
883e09c
 
 
f7b9ae7
883e09c
f7b9ae7
883e09c
f7b9ae7
 
883e09c
 
 
f7b9ae7
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
---
library_name: peft
base_model: microsoft/Phi-3-mini-4k-instruct
tags:
  - code-review
  - qlora
  - lora
  - fine-tuned
  - code-analysis
  - phi-3
license: mit
language:
  - en
pipeline_tag: text-generation
---

# Model Card for phi3-mini-code-reviewer

A QLoRA fine-tuned version of `microsoft/Phi-3-mini-4k-instruct`, specialised to review short
Python functions and return a **structured JSON code review** β€” issues by category and severity,
actionable fix suggestions, and an overall approve/request-changes verdict.

## Model Details

### Model Description

This model takes a Python function as input and returns a strict JSON review object, similar to
a first-pass automated code reviewer. It was fine-tuned to close the gap between a general-purpose
instruction model's inconsistent, prose-heavy code commentary and a schema-conformant, structured
review a review-automation pipeline can actually parse and act on.

- **Developed by:** Themal De Silva
- **Model type:** Causal decoder-only LLM, LoRA-adapted (merged)
- **Language(s):** English (input/output), Python (code domain)
- **License:** MIT (inherited from base model)
- **Finetuned from model:** [microsoft/Phi-3-mini-4k-instruct](https://huggingface.co/microsoft/Phi-3-mini-4k-instruct)

### Model Sources

- **Repository:** CDAZZDEV-MLE-Themal/task2_genai (see notebook `task2_finetuning.ipynb`)

## Uses

### Direct Use

Given a Python function (roughly 15–45 lines) as the user turn, the model returns a JSON object:

```json
{
  "issues": [
    {"category": "bug|security|performance|style|readability",
     "severity": "critical|major|minor",
     "line_hint": "<short quote or line description>",
     "suggestion": "<specific, actionable fix>"}
  ],
  "overall_verdict": "approve|request_changes",
  "summary": "<2-3 sentence summary>"
}
```

Intended as a first-pass automated reviewer to flag likely issues for a human reviewer to confirm
β€” not a replacement for human code review.

### Out-of-Scope Use

- Not evaluated on languages other than Python, or on files longer than ~45 lines / outside a
  4096-token context.
- Not a security-audit tool: manual review found the model under-detects security issues
  relative to bug/style issues (see Evaluation below) β€” do not rely on it as a sole security gate.
- Not intended for general-purpose chat; it was trained exclusively on the code-review task and
  its outputs outside that format are unvalidated.

## Bias, Risks, and Limitations

- **Schema drift:** in manual testing, most outputs used categories/severities close to but not
  strictly matching the intended enum (e.g. `"Medium"` instead of `"major"`) β€” downstream
  consumers should validate/normalise the output rather than assume strict enum compliance.
- **Under-detection of security issues:** the training data (100+ teacher-generated examples)
  under-represented security-critical scenarios relative to bugs/style; the model is more likely
  to miss a real vulnerability than to hallucinate one, but it does miss some (e.g. failed to
  flag an `eval()` injection vulnerability in one held-out test case).
- **Small fine-tuning set:** trained on ~85 examples (90 train / 10 val / 10 test split from ~105
  generated), which limits generalisation to code patterns outside the ~20 scenario types used
  for data generation (see Training Data below).
- **Occasional hallucination:** manual review of 10 held-out outputs found 1 hallucinated issue
  (an invented stack-overflow concern in code with no recursion), a 10% rate in that sample.

### Recommendations

Treat outputs as a first-pass triage signal, always paired with human review, especially for
security-sensitive code. Validate/coerce the returned category and severity fields against the
intended enum before using them programmatically.

## How to Get Started with the Model

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "Themal/phi3-mini-code-reviewer"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")

system_prompt = (
    "You are an automated Python code reviewer. Given a code snippet, respond with a "
    "single strict JSON object: issues (category, severity, line_hint, suggestion), "
    "overall_verdict, and summary. No text outside the JSON."
)
code_snippet = '''def divide(a, b):
    return a / b
'''

chat = [{"role": "system", "content": system_prompt},
        {"role": "user", "content": code_snippet}]
prompt = tokenizer.apply_chat_template(chat, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=400, do_sample=False, pad_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```

## Training Details

### Training Data

~105 synthetic (code, review) pairs generated by `openai/gpt-oss-120b` (teacher model, via Groq)
across 20 hand-written scenario seeds (Flask endpoints, pandas pipelines, retry wrappers, JWT
auth, CSV parsing, thread pools, etc.) crossed with 5 issue-mix instructions, each example
containing 1–3 deliberately planted realistic issues. Diversity was checked via prompt-length
distribution, issue-category frequency, and scenario coverage before training. Split 80/10/10
into train/validation/test.

### Training Procedure

QLoRA fine-tuning: base model loaded in 4-bit NF4 quantization (bitsandbytes, double quant, bf16
compute dtype), LoRA adapters applied to all attention and MLP projection layers, trained for 3
epochs, then merged into the base model at full (bf16) precision post-training (adapters were
merged onto a freshly reloaded full-precision copy of the base model rather than the 4-bit
training copy, to avoid known merge instability with quantized layers).

#### Preprocessing

Examples formatted using the base model's native chat template
(`<|system|>...<|user|>...<|assistant|>...`), with the assistant turn set to the reference
review's JSON serialised as a string.

#### Training Hyperparameters

- **Training regime:** bf16 compute dtype, 4-bit NF4 quantized base weights during training
- **LoRA rank (r):** 16
- **LoRA alpha:** 32
- **LoRA dropout:** 0.05
- **Target modules:** q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
- **Learning rate:** 2e-4, cosine schedule, 3% warmup
- **Epochs:** 3
- **Batch size:** 2 (per device), gradient accumulation 8 (effective batch size 16)
- **Max sequence length:** 1024 tokens

#### Speeds, Sizes, Times

- **Hardware:** single Google Colab T4 GPU (free tier)
- **Trainable parameters:** 8,912,896 / 3,829,992,448 total (0.23%)

## Evaluation

### Testing Data, Factors & Metrics

#### Testing Data

10 held-out examples from the same generation process as training data (never seen during
training or validation).

#### Metrics

ROUGE-L (F1), BERTScore (F1), and LLM-as-judge (`openai/gpt-oss-120b`) scoring issue_detection,
json_validity, and actionability on a 1–5 scale, plus a manual hallucination review of 10
fine-tuned outputs labelled correct/partial/hallucinated.

### Results

| Metric | Base (Phi-3-mini, no fine-tuning) | Fine-tuned |
|---|---|---|
| BERTScore F1 | 0.884 | 0.888 |
| LLM-judge: issue_detection (1-5) | 1.30 | 1.80 |
| LLM-judge: json_validity (1-5) | 4.40 | 4.70 |
| LLM-judge: actionability (1-5) | 2.60 | 2.90 |

Manual review of 10 fine-tuned outputs: 6 correct, 3 partial, 1 hallucinated (10% hallucination
rate).

#### Summary

Fine-tuning improved every measured dimension, most notably issue_detection (+0.5) and
actionability (+0.3). The main remaining gap is schema conformance β€” outputs are valid JSON but
frequently drift from the intended category/severity enum β€” and under-detection of
security-critical issues specifically, traced to under-representation of security scenarios in
the training data. See the full evaluation notebook for per-example detail.

## Environmental Impact

- **Hardware Type:** NVIDIA T4 (Google Colab free tier)
- **Hours used:** < 1 hour (QLoRA fine-tuning, 3 epochs, ~85 training examples)
- **Cloud Provider:** Google Cloud (via Colab)
- **Compute Region:** Unknown (Colab-assigned)
- **Carbon Emitted:** Not measured; given the short training time and single T4, expected to be
  minimal relative to full fine-tuning or larger models.

## Technical Specifications

### Model Architecture and Objective

Decoder-only transformer (Phi-3-mini architecture, 3.8B parameters), causal language modeling
objective, adapted via low-rank (LoRA) weight updates on attention and MLP projections, merged
into the base weights post-training. Objective during fine-tuning: supervised next-token
prediction on (code, structured-JSON-review) chat-formatted pairs.

### Compute Infrastructure

#### Hardware

Single NVIDIA T4 GPU, Google Colab free tier.

#### Software

`transformers`, `peft`, `bitsandbytes` (4-bit NF4 quantization), `trl` (SFTTrainer), `datasets`.

## Citation

This model was produced as part of a technical assessment (Ceylon Dazzling Dev Holding Senior
MLE Assessment, Task 2). No formal publication.

## Model Card Contact

Themal De Silva