Instructions to use PrithviRana/DevOps with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use PrithviRana/DevOps with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf PrithviRana/DevOps:Q4_K_M # Run inference directly in the terminal: llama cli -hf PrithviRana/DevOps:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf PrithviRana/DevOps:Q4_K_M # Run inference directly in the terminal: llama cli -hf PrithviRana/DevOps:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf PrithviRana/DevOps:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf PrithviRana/DevOps:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf PrithviRana/DevOps:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf PrithviRana/DevOps:Q4_K_M
Use Docker
docker model run hf.co/PrithviRana/DevOps:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use PrithviRana/DevOps with Ollama:
ollama run hf.co/PrithviRana/DevOps:Q4_K_M
- Unsloth Desktop
- Pi
How to use PrithviRana/DevOps with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PrithviRana/DevOps:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "PrithviRana/DevOps:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use PrithviRana/DevOps with Docker Model Runner:
docker model run hf.co/PrithviRana/DevOps:Q4_K_M
- Lemonade
How to use PrithviRana/DevOps with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull PrithviRana/DevOps:Q4_K_M
Run and chat with the model
lemonade run user.DevOps-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use PrithviRana/DevOps with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PrithviRana/DevOps:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default PrithviRana/DevOps:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use PrithviRana/DevOps with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PrithviRana/DevOps:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "PrithviRana/DevOps:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download README.md from PrithviRana/DevOps: direct link, hf CLI and curl.
- Browser
- Download file 9.13 kB
-
https://huggingface.co/PrithviRana/DevOps/resolve/main/README.md
- Command line
-
hf download hf://PrithviRana/DevOps/README.md
-
curl -L -o README.md https://huggingface.co/PrithviRana/DevOps/resolve/main/README.md
9.13 kB
| # DevOps Qwen — Fine-Tuned Qwen2.5-3B-Instruct | |
| A Qwen2.5-3B-Instruct model fine-tuned with **LoRA (Low-Rank Adaptation)** on a DevOps-focused dataset. | |
| The model is designed for practical DevOps, Cloud, Linux, Docker, Kubernetes, Terraform, CI/CD, networking, monitoring, and troubleshooting questions. | |
| ## Model Details | |
| | Property | Value | | |
| | ------------------- | ------------------------------ | | |
| | Base Model | Qwen/Qwen2.5-3B-Instruct | | |
| | Model Type | Causal Language Model | | |
| | Fine-Tuning | LoRA | | |
| | LoRA Rank | 8 | | |
| | LoRA Alpha | 16 | | |
| | LoRA Dropout | 0.05 | | |
| | Target Modules | q_proj, k_proj, v_proj, o_proj | | |
| | Training Epochs | 1 | | |
| | Max Sequence Length | 512 | | |
| | Quantization | Q4_K_M | | |
| | Format | GGUF | | |
| | Approx. Model Size | 1.9 GB | | |
| | Runtime | Ollama / llama.cpp | | |
| | Primary Use | DevOps AI Assistant | | |
| ## What is this model? | |
| This model is a specialized version of Qwen2.5-3B-Instruct trained on DevOps-oriented examples. | |
| The goal of the fine-tuning is to improve the model's ability to provide practical responses for: | |
| * Linux administration | |
| * AWS | |
| * GCP | |
| * Docker | |
| * Kubernetes | |
| * Terraform | |
| * Git | |
| * Jenkins | |
| * CI/CD | |
| * Networking | |
| * HTTP troubleshooting | |
| * Monitoring | |
| * Production troubleshooting | |
| * Cloud infrastructure | |
| The model is intended to provide answers with: | |
| 1. Root cause or explanation | |
| 2. Exact commands where appropriate | |
| 3. Short explanation of commands | |
| 4. Production-safe troubleshooting steps | |
| ## Fine-Tuning Approach | |
| The model was fine-tuned using **LoRA — Low-Rank Adaptation**. | |
| Instead of updating the entire base model, LoRA trains a small set of additional parameters while keeping most of the original model frozen. | |
| ```text | |
| Qwen2.5-3B-Instruct | |
| | | |
| v | |
| LoRA Training | |
| | | |
| v | |
| LoRA Adapter | |
| | | |
| v | |
| Merge Adapter + Base Model | |
| | | |
| v | |
| Merged Model | |
| ``` | |
| This approach reduces training memory and computational requirements compared with full fine-tuning. | |
| ## Training Configuration | |
| ```text | |
| Base Model: | |
| Qwen/Qwen2.5-3B-Instruct | |
| LoRA: | |
| r = 8 | |
| alpha = 16 | |
| dropout = 0.05 | |
| Target modules: | |
| q_proj | |
| k_proj | |
| v_proj | |
| o_proj | |
| Epochs: | |
| 1 | |
| Batch size: | |
| 1 | |
| Gradient accumulation: | |
| 4 | |
| Learning rate: | |
| 2e-4 | |
| Maximum sequence length: | |
| 512 | |
| ``` | |
| ## Model Conversion | |
| After LoRA training, the adapter was merged with the base model. | |
| The merged Hugging Face model was then converted to GGUF using llama.cpp. | |
| ```text | |
| LoRA Adapter | |
| | | |
| v | |
| Merged Hugging Face Model | |
| | | |
| v | |
| GGUF F16 | |
| | | |
| v | |
| Q4_K_M Quantization | |
| | | |
| v | |
| qwen-devops-q4_k_m.gguf | |
| ``` | |
| ### GGUF | |
| **GGUF (GPT-Generated Unified Format)** is an efficient model format commonly used for local LLM inference with llama.cpp and compatible runtimes. | |
| ### Q4_K_M | |
| Q4_K_M is a 4-bit quantization format. | |
| It reduces model storage and memory requirements while maintaining a useful level of model quality for local inference. | |
| Approximate sizes: | |
| ```text | |
| Merged Hugging Face Model ≈ 12 GB | |
| GGUF F16 ≈ 5.8 GB | |
| Q4_K_M GGUF ≈ 1.9 GB | |
| ``` | |
| ## Hardware Used | |
| The model was developed and tested in a CPU-only environment. | |
| ```text | |
| CPU: | |
| AMD EPYC 7543 | |
| CPU cores available: | |
| 12 | |
| RAM: | |
| ~57 GB | |
| GPU: | |
| None | |
| CUDA: | |
| False | |
| Python: | |
| 3.10.14 | |
| ``` | |
| ## Usage with Ollama | |
| Download the GGUF model from this repository. | |
| Create a `Modelfile`: | |
| ```text | |
| FROM ./qwen-devops-q4_k_m.gguf | |
| PARAMETER temperature 0.2 | |
| PARAMETER top_k 20 | |
| PARAMETER top_p 0.9 | |
| PARAMETER repeat_penalty 1.1 | |
| PARAMETER num_ctx 4096 | |
| SYSTEM """ | |
| You are a senior DevOps and Cloud engineer. | |
| Give practical and accurate technical answers. | |
| For Linux, AWS, Docker, Kubernetes, Terraform, Git, | |
| Jenkins, CI/CD, monitoring, networking and troubleshooting: | |
| - Explain the root cause. | |
| - Give exact commands when appropriate. | |
| - Explain commands briefly. | |
| - Do not invent information. | |
| - If you don't know something, clearly say so. | |
| - Prefer safe production-ready solutions. | |
| """ | |
| ``` | |
| Create the Ollama model: | |
| ```bash | |
| ollama create devops-qwen -f Modelfile | |
| ``` | |
| Run: | |
| ```bash | |
| ollama run devops-qwen | |
| ``` | |
| ## Example | |
| Question: | |
| ```text | |
| How do I troubleshoot a 502 Bad Gateway error from an AWS ALB? | |
| ``` | |
| The model is intended to provide a structured troubleshooting approach such as: | |
| ```text | |
| 1. Check ALB target health | |
| 2. Verify application is listening on the expected port | |
| 3. Check security groups | |
| 4. Check target response | |
| 5. Review ALB access logs | |
| 6. Review application logs | |
| 7. Test the target directly | |
| 8. Check health-check configuration | |
| ``` | |
| Example commands may include: | |
| ```bash | |
| ss -lntp | |
| curl -v http://127.0.0.1:8080/ | |
| curl -v http://TARGET_PRIVATE_IP:8080/ | |
| ``` | |
| ## Ollama API | |
| Non-streaming request: | |
| ```bash | |
| curl http://localhost:11434/api/generate \ | |
| -d '{ | |
| "model": "devops-qwen", | |
| "prompt": "How do I check disk usage in Linux?", | |
| "stream": false | |
| }' | |
| ``` | |
| Streaming request: | |
| ```bash | |
| curl http://localhost:11434/api/generate \ | |
| -d '{ | |
| "model": "devops-qwen", | |
| "prompt": "How do I troubleshoot Kubernetes CrashLoopBackOff?", | |
| "stream": true | |
| }' | |
| ``` | |
| ## llama.cpp | |
| The GGUF model can also be used with llama.cpp: | |
| ```bash | |
| ./llama-cli \ | |
| -m qwen-devops-q4_k_m.gguf | |
| ``` | |
| ## Recommended Generation Parameters | |
| For technical and DevOps questions: | |
| ```text | |
| temperature = 0.2 | |
| top_k = 20 | |
| top_p = 0.9 | |
| repeat_penalty = 1.1 | |
| context = 4096 | |
| ``` | |
| Lower temperature is used to encourage more deterministic and consistent technical responses. | |
| ## Fine-Tuning vs RAG | |
| This model should not be considered a replacement for RAG. | |
| Fine-tuning is useful for: | |
| * Response style | |
| * Domain behavior | |
| * Task patterns | |
| * DevOps troubleshooting patterns | |
| * Command-oriented responses | |
| RAG is useful for: | |
| * Company documentation | |
| * Current infrastructure information | |
| * Internal runbooks | |
| * AWS architecture documentation | |
| * Frequently changing configuration | |
| * Private knowledge bases | |
| Recommended architecture: | |
| ```text | |
| User | |
| | | |
| v | |
| Chat UI | |
| | | |
| v | |
| n8n / FastAPI | |
| | | |
| v | |
| RAG Retriever | |
| | | |
| v | |
| Vector Database | |
| | | |
| v | |
| Relevant DevOps Documents | |
| | | |
| v | |
| Context | |
| | | |
| v | |
| devops-qwen | |
| | | |
| v | |
| Final Answer | |
| ``` | |
| ## Intended Use | |
| This model is intended for: | |
| * DevOps assistants | |
| * Cloud troubleshooting assistants | |
| * Linux support | |
| * Infrastructure automation | |
| * CI/CD assistance | |
| * Kubernetes troubleshooting | |
| * Terraform assistance | |
| * Internal technical assistants | |
| * RAG-based DevOps assistants | |
| ## Limitations | |
| The model is relatively small at approximately 3B parameters. | |
| It may: | |
| * Make incorrect technical assumptions | |
| * Produce outdated information | |
| * Generate commands that require environment-specific changes | |
| * Fail on complex infrastructure architecture | |
| * Require RAG or external tools for current infrastructure information | |
| Always verify commands before running them in production. | |
| For production environments, use appropriate: | |
| * Backups | |
| * Change management | |
| * Testing | |
| * Access controls | |
| * Approval processes | |
| ## Security | |
| Do not provide the model with: | |
| * AWS access keys | |
| * Private SSH keys | |
| * Passwords | |
| * API tokens | |
| * Database credentials | |
| * TLS private keys | |
| * Other secrets | |
| When integrating this model with automation, use least-privilege credentials and approval controls for destructive operations. | |
| ## Project Pipeline | |
| ```text | |
| DevOps Dataset | |
| | | |
| v | |
| JSONL Validation | |
| | | |
| v | |
| Train / Validation Split | |
| | | |
| v | |
| Qwen2.5-3B-Instruct | |
| | | |
| v | |
| LoRA Fine-Tuning | |
| | | |
| v | |
| LoRA Adapter | |
| | | |
| v | |
| Merge | |
| | | |
| v | |
| Merged Model | |
| | | |
| v | |
| GGUF F16 | |
| | | |
| v | |
| Q4_K_M | |
| | | |
| v | |
| Ollama | |
| | | |
| v | |
| devops-qwen | |
| | | |
| v | |
| API / n8n / RAG | |
| ``` | |
| ## Benchmark | |
| The project includes an automated benchmark comparing: | |
| ```text | |
| qwen2.5:3b | |
| VS | |
| devops-qwen | |
| ``` | |
| The benchmark contains 10 DevOps questions covering: | |
| * Linux | |
| * AWS ALB | |
| * Docker | |
| * CPU/RAM | |
| * Disk usage | |
| * Terraform | |
| * Kubernetes | |
| * HTTP | |
| * Production troubleshooting | |
| Benchmark output: | |
| ```text | |
| benchmark_results.json | |
| ``` | |
| ## Model Card Summary | |
| ```text | |
| Model: | |
| DevOps Qwen | |
| Base: | |
| Qwen2.5-3B-Instruct | |
| Fine-Tuning: | |
| LoRA | |
| Format: | |
| GGUF | |
| Quantization: | |
| Q4_K_M | |
| Size: | |
| ~1.9 GB | |
| Runtime: | |
| Ollama / llama.cpp | |
| Domain: | |
| DevOps / Cloud / Infrastructure | |
| Recommended: | |
| CPU local inference + RAG | |
| ``` | |
| ## License | |
| This model is derived from Qwen2.5-3B-Instruct. | |
| Users should review and comply with the applicable Qwen model license and its terms before using or redistributing this model, particularly for commercial use. | |
| The fine-tuning dataset and any additional project components may have their own applicable terms. | |
| ## Disclaimer | |
| This model is an experimental DevOps-focused AI assistant. It is not a substitute for production change-control procedures or expert review. | |
| Always validate generated commands and infrastructure changes before applying them to production systems. | |