File size: 2,613 Bytes
f8ff5f4
d62e46d
f8ff5f4
 
d62e46d
 
 
 
 
 
 
f8ff5f4
d62e46d
 
f8ff5f4
 
9bcaa48
d62e46d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f8ff5f4
d62e46d
f8ff5f4
d62e46d
 
 
 
9bcaa48
d62e46d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f8ff5f4
9bcaa48
 
d62e46d
9bcaa48
d62e46d
9bcaa48
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
---
license: apache-2.0
base_model: LiquidAI/LFM2.5-1.2B-Thinking
tags:
  - lora
  - unsloth
  - reasoning
  - distillation
  - lfm2
datasets:
  - TeichAI/gpt-5.2-high-reasoning-250x
language:
  - en
pipeline_tag: text-generation
---

# Micro-Merlin-Experimental

This is a fine-tune of **LiquidAI/LFM2.5-1.2B-Thinking** on GPT-5.2 reasoning traces.

The model was trained with LoRA on the [TeichAI/gpt-5.2-high-reasoning-250x](https://huggingface.co/datasets/TeichAI/gpt-5.2-high-reasoning-250x) dataset, a collection of high-reasoning-depth traces distilled from GPT-5.2, focused on production-grade DevOps, backend, and infrastructure engineering tasks. The goal is to transfer GPT-5.2's structured `<think>` reasoning style onto a compact 1.2B model that runs comfortably on consumer hardware.

## Model & Training Details

| Field | Value |
|---|---|
| **Base model** | LiquidAI/LFM2.5-1.2B-Thinking |
| **Parameters** | 1.2B |
| **Method** | LoRA (16-bit, rank-stabilized) |
| **Dataset** | TeichAI/gpt-5.2-high-reasoning-250x |
| **Training examples** | 249 |
| **Epochs** | 1 |
| **Total steps** | ~63 |
| **Final training loss** | 2.121 |
| **LoRA rank (r)** | 64 |
| **LoRA alpha** | 64 |
| **LoRA dropout** | 0 |
| **rsLoRA** | Enabled |
| **Target modules** | q_proj, k_proj, v_proj, out_proj, in_proj, w1, w2, w3 |
| **Max sequence length** | 20,480 |
| **Batch size (effective)** | 4 (1 × 4 grad. accum.) |
| **Learning rate** | 2e-4 |
| **LR scheduler** | Cosine |
| **Warmup steps** | 3 |
| **Optimizer** | adamw_8bit |
| **Weight decay** | 0.01 |
| **Precision** | FP16 |
| **Loss masking** | Responses only (`<think>` + answer) |
| **Hardware** | 1× NVIDIA Tesla T4 (16 GB) |
| **Framework** | Unsloth + TRL SFTTrainer |
| **Training runtime** | ~608 s (~10 min) |
| **Chat template** | ChatML (`<|im_start|>` / `<|im_end|>`) |

## Usage

```python
from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="OrionLLM/Micro-Merlin-Experimental",
    max_seq_length=20480,
    load_in_4bit=False,
)
FastLanguageModel.for_inference(model)

messages = [{"role": "user", "content": "Design a rate limiter for a REST API."}]
inputs = tokenizer.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True, return_tensors="pt"
).to(model.device)

out = model.generate(**inputs, max_new_tokens=1024, temperature=0.5, repetition_penalty=1.15)
print(tokenizer.decode(out[0], skip_special_tokens=True))
```

---

<div align="center">

**Merlin Research • 2026**

Developed by [DedeProGames](https://huggingface.co/DedeProGames)

</div>