PrithviRana commited on
Commit
5038ec4
·
verified ·
1 Parent(s): 1137dd5

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +510 -0
README.md ADDED
@@ -0,0 +1,510 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # DevOps Qwen — Fine-Tuned Qwen2.5-3B-Instruct
2
+
3
+ A Qwen2.5-3B-Instruct model fine-tuned with **LoRA (Low-Rank Adaptation)** on a DevOps-focused dataset.
4
+
5
+ The model is designed for practical DevOps, Cloud, Linux, Docker, Kubernetes, Terraform, CI/CD, networking, monitoring, and troubleshooting questions.
6
+
7
+ ## Model Details
8
+
9
+ | Property | Value |
10
+ | ------------------- | ------------------------------ |
11
+ | Base Model | Qwen/Qwen2.5-3B-Instruct |
12
+ | Model Type | Causal Language Model |
13
+ | Fine-Tuning | LoRA |
14
+ | LoRA Rank | 8 |
15
+ | LoRA Alpha | 16 |
16
+ | LoRA Dropout | 0.05 |
17
+ | Target Modules | q_proj, k_proj, v_proj, o_proj |
18
+ | Training Epochs | 1 |
19
+ | Max Sequence Length | 512 |
20
+ | Quantization | Q4_K_M |
21
+ | Format | GGUF |
22
+ | Approx. Model Size | 1.9 GB |
23
+ | Runtime | Ollama / llama.cpp |
24
+ | Primary Use | DevOps AI Assistant |
25
+
26
+ ## What is this model?
27
+
28
+ This model is a specialized version of Qwen2.5-3B-Instruct trained on DevOps-oriented examples.
29
+
30
+ The goal of the fine-tuning is to improve the model's ability to provide practical responses for:
31
+
32
+ * Linux administration
33
+ * AWS
34
+ * GCP
35
+ * Docker
36
+ * Kubernetes
37
+ * Terraform
38
+ * Git
39
+ * Jenkins
40
+ * CI/CD
41
+ * Networking
42
+ * HTTP troubleshooting
43
+ * Monitoring
44
+ * Production troubleshooting
45
+ * Cloud infrastructure
46
+
47
+ The model is intended to provide answers with:
48
+
49
+ 1. Root cause or explanation
50
+ 2. Exact commands where appropriate
51
+ 3. Short explanation of commands
52
+ 4. Production-safe troubleshooting steps
53
+
54
+ ## Fine-Tuning Approach
55
+
56
+ The model was fine-tuned using **LoRA — Low-Rank Adaptation**.
57
+
58
+ Instead of updating the entire base model, LoRA trains a small set of additional parameters while keeping most of the original model frozen.
59
+
60
+ ```text
61
+ Qwen2.5-3B-Instruct
62
+ |
63
+ v
64
+ LoRA Training
65
+ |
66
+ v
67
+ LoRA Adapter
68
+ |
69
+ v
70
+ Merge Adapter + Base Model
71
+ |
72
+ v
73
+ Merged Model
74
+ ```
75
+
76
+ This approach reduces training memory and computational requirements compared with full fine-tuning.
77
+
78
+ ## Training Configuration
79
+
80
+ ```text
81
+ Base Model:
82
+ Qwen/Qwen2.5-3B-Instruct
83
+
84
+ LoRA:
85
+ r = 8
86
+ alpha = 16
87
+ dropout = 0.05
88
+
89
+ Target modules:
90
+ q_proj
91
+ k_proj
92
+ v_proj
93
+ o_proj
94
+
95
+ Epochs:
96
+ 1
97
+
98
+ Batch size:
99
+ 1
100
+
101
+ Gradient accumulation:
102
+ 4
103
+
104
+ Learning rate:
105
+ 2e-4
106
+
107
+ Maximum sequence length:
108
+ 512
109
+ ```
110
+
111
+ ## Model Conversion
112
+
113
+ After LoRA training, the adapter was merged with the base model.
114
+
115
+ The merged Hugging Face model was then converted to GGUF using llama.cpp.
116
+
117
+ ```text
118
+ LoRA Adapter
119
+ |
120
+ v
121
+ Merged Hugging Face Model
122
+ |
123
+ v
124
+ GGUF F16
125
+ |
126
+ v
127
+ Q4_K_M Quantization
128
+ |
129
+ v
130
+ qwen-devops-q4_k_m.gguf
131
+ ```
132
+
133
+ ### GGUF
134
+
135
+ **GGUF (GPT-Generated Unified Format)** is an efficient model format commonly used for local LLM inference with llama.cpp and compatible runtimes.
136
+
137
+ ### Q4_K_M
138
+
139
+ Q4_K_M is a 4-bit quantization format.
140
+
141
+ It reduces model storage and memory requirements while maintaining a useful level of model quality for local inference.
142
+
143
+ Approximate sizes:
144
+
145
+ ```text
146
+ Merged Hugging Face Model ≈ 12 GB
147
+ GGUF F16 ≈ 5.8 GB
148
+ Q4_K_M GGUF ≈ 1.9 GB
149
+ ```
150
+
151
+ ## Hardware Used
152
+
153
+ The model was developed and tested in a CPU-only environment.
154
+
155
+ ```text
156
+ CPU:
157
+ AMD EPYC 7543
158
+
159
+ CPU cores available:
160
+ 12
161
+
162
+ RAM:
163
+ ~57 GB
164
+
165
+ GPU:
166
+ None
167
+
168
+ CUDA:
169
+ False
170
+
171
+ Python:
172
+ 3.10.14
173
+ ```
174
+
175
+ ## Usage with Ollama
176
+
177
+ Download the GGUF model from this repository.
178
+
179
+ Create a `Modelfile`:
180
+
181
+ ```text
182
+ FROM ./qwen-devops-q4_k_m.gguf
183
+
184
+ PARAMETER temperature 0.2
185
+ PARAMETER top_k 20
186
+ PARAMETER top_p 0.9
187
+ PARAMETER repeat_penalty 1.1
188
+ PARAMETER num_ctx 4096
189
+
190
+ SYSTEM """
191
+ You are a senior DevOps and Cloud engineer.
192
+
193
+ Give practical and accurate technical answers.
194
+
195
+ For Linux, AWS, Docker, Kubernetes, Terraform, Git,
196
+ Jenkins, CI/CD, monitoring, networking and troubleshooting:
197
+
198
+ - Explain the root cause.
199
+ - Give exact commands when appropriate.
200
+ - Explain commands briefly.
201
+ - Do not invent information.
202
+ - If you don't know something, clearly say so.
203
+ - Prefer safe production-ready solutions.
204
+ """
205
+ ```
206
+
207
+ Create the Ollama model:
208
+
209
+ ```bash
210
+ ollama create devops-qwen -f Modelfile
211
+ ```
212
+
213
+ Run:
214
+
215
+ ```bash
216
+ ollama run devops-qwen
217
+ ```
218
+
219
+ ## Example
220
+
221
+ Question:
222
+
223
+ ```text
224
+ How do I troubleshoot a 502 Bad Gateway error from an AWS ALB?
225
+ ```
226
+
227
+ The model is intended to provide a structured troubleshooting approach such as:
228
+
229
+ ```text
230
+ 1. Check ALB target health
231
+ 2. Verify application is listening on the expected port
232
+ 3. Check security groups
233
+ 4. Check target response
234
+ 5. Review ALB access logs
235
+ 6. Review application logs
236
+ 7. Test the target directly
237
+ 8. Check health-check configuration
238
+ ```
239
+
240
+ Example commands may include:
241
+
242
+ ```bash
243
+ ss -lntp
244
+ curl -v http://127.0.0.1:8080/
245
+ curl -v http://TARGET_PRIVATE_IP:8080/
246
+ ```
247
+
248
+ ## Ollama API
249
+
250
+ Non-streaming request:
251
+
252
+ ```bash
253
+ curl http://localhost:11434/api/generate \
254
+ -d '{
255
+ "model": "devops-qwen",
256
+ "prompt": "How do I check disk usage in Linux?",
257
+ "stream": false
258
+ }'
259
+ ```
260
+
261
+ Streaming request:
262
+
263
+ ```bash
264
+ curl http://localhost:11434/api/generate \
265
+ -d '{
266
+ "model": "devops-qwen",
267
+ "prompt": "How do I troubleshoot Kubernetes CrashLoopBackOff?",
268
+ "stream": true
269
+ }'
270
+ ```
271
+
272
+ ## llama.cpp
273
+
274
+ The GGUF model can also be used with llama.cpp:
275
+
276
+ ```bash
277
+ ./llama-cli \
278
+ -m qwen-devops-q4_k_m.gguf
279
+ ```
280
+
281
+ ## Recommended Generation Parameters
282
+
283
+ For technical and DevOps questions:
284
+
285
+ ```text
286
+ temperature = 0.2
287
+ top_k = 20
288
+ top_p = 0.9
289
+ repeat_penalty = 1.1
290
+ context = 4096
291
+ ```
292
+
293
+ Lower temperature is used to encourage more deterministic and consistent technical responses.
294
+
295
+ ## Fine-Tuning vs RAG
296
+
297
+ This model should not be considered a replacement for RAG.
298
+
299
+ Fine-tuning is useful for:
300
+
301
+ * Response style
302
+ * Domain behavior
303
+ * Task patterns
304
+ * DevOps troubleshooting patterns
305
+ * Command-oriented responses
306
+
307
+ RAG is useful for:
308
+
309
+ * Company documentation
310
+ * Current infrastructure information
311
+ * Internal runbooks
312
+ * AWS architecture documentation
313
+ * Frequently changing configuration
314
+ * Private knowledge bases
315
+
316
+ Recommended architecture:
317
+
318
+ ```text
319
+ User
320
+ |
321
+ v
322
+ Chat UI
323
+ |
324
+ v
325
+ n8n / FastAPI
326
+ |
327
+ v
328
+ RAG Retriever
329
+ |
330
+ v
331
+ Vector Database
332
+ |
333
+ v
334
+ Relevant DevOps Documents
335
+ |
336
+ v
337
+ Context
338
+ |
339
+ v
340
+ devops-qwen
341
+ |
342
+ v
343
+ Final Answer
344
+ ```
345
+
346
+ ## Intended Use
347
+
348
+ This model is intended for:
349
+
350
+ * DevOps assistants
351
+ * Cloud troubleshooting assistants
352
+ * Linux support
353
+ * Infrastructure automation
354
+ * CI/CD assistance
355
+ * Kubernetes troubleshooting
356
+ * Terraform assistance
357
+ * Internal technical assistants
358
+ * RAG-based DevOps assistants
359
+
360
+ ## Limitations
361
+
362
+ The model is relatively small at approximately 3B parameters.
363
+
364
+ It may:
365
+
366
+ * Make incorrect technical assumptions
367
+ * Produce outdated information
368
+ * Generate commands that require environment-specific changes
369
+ * Fail on complex infrastructure architecture
370
+ * Require RAG or external tools for current infrastructure information
371
+
372
+ Always verify commands before running them in production.
373
+
374
+ For production environments, use appropriate:
375
+
376
+ * Backups
377
+ * Change management
378
+ * Testing
379
+ * Access controls
380
+ * Approval processes
381
+
382
+ ## Security
383
+
384
+ Do not provide the model with:
385
+
386
+ * AWS access keys
387
+ * Private SSH keys
388
+ * Passwords
389
+ * API tokens
390
+ * Database credentials
391
+ * TLS private keys
392
+ * Other secrets
393
+
394
+ When integrating this model with automation, use least-privilege credentials and approval controls for destructive operations.
395
+
396
+ ## Project Pipeline
397
+
398
+ ```text
399
+ DevOps Dataset
400
+ |
401
+ v
402
+ JSONL Validation
403
+ |
404
+ v
405
+ Train / Validation Split
406
+ |
407
+ v
408
+ Qwen2.5-3B-Instruct
409
+ |
410
+ v
411
+ LoRA Fine-Tuning
412
+ |
413
+ v
414
+ LoRA Adapter
415
+ |
416
+ v
417
+ Merge
418
+ |
419
+ v
420
+ Merged Model
421
+ |
422
+ v
423
+ GGUF F16
424
+ |
425
+ v
426
+ Q4_K_M
427
+ |
428
+ v
429
+ Ollama
430
+ |
431
+ v
432
+ devops-qwen
433
+ |
434
+ v
435
+ API / n8n / RAG
436
+ ```
437
+
438
+ ## Benchmark
439
+
440
+ The project includes an automated benchmark comparing:
441
+
442
+ ```text
443
+ qwen2.5:3b
444
+ VS
445
+ devops-qwen
446
+ ```
447
+
448
+ The benchmark contains 10 DevOps questions covering:
449
+
450
+ * Linux
451
+ * AWS ALB
452
+ * Docker
453
+ * CPU/RAM
454
+ * Disk usage
455
+ * Terraform
456
+ * Kubernetes
457
+ * HTTP
458
+ * Production troubleshooting
459
+
460
+ Benchmark output:
461
+
462
+ ```text
463
+ benchmark_results.json
464
+ ```
465
+
466
+ ## Model Card Summary
467
+
468
+ ```text
469
+ Model:
470
+ DevOps Qwen
471
+
472
+ Base:
473
+ Qwen2.5-3B-Instruct
474
+
475
+ Fine-Tuning:
476
+ LoRA
477
+
478
+ Format:
479
+ GGUF
480
+
481
+ Quantization:
482
+ Q4_K_M
483
+
484
+ Size:
485
+ ~1.9 GB
486
+
487
+ Runtime:
488
+ Ollama / llama.cpp
489
+
490
+ Domain:
491
+ DevOps / Cloud / Infrastructure
492
+
493
+ Recommended:
494
+ CPU local inference + RAG
495
+ ```
496
+
497
+ ## License
498
+
499
+ This model is derived from Qwen2.5-3B-Instruct.
500
+
501
+ Users should review and comply with the applicable Qwen model license and its terms before using or redistributing this model, particularly for commercial use.
502
+
503
+ The fine-tuning dataset and any additional project components may have their own applicable terms.
504
+
505
+ ## Disclaimer
506
+
507
+ This model is an experimental DevOps-focused AI assistant. It is not a substitute for production change-control procedures or expert review.
508
+
509
+ Always validate generated commands and infrastructure changes before applying them to production systems.
510
+