DKube-instruct-v2-7B

A LoRA adapter on Qwen/Qwen2.5-7B-Instruct for Kubernetes troubleshooting. Given an error (ImagePullBackOff, CrashLoopBackOff, exit codes, scheduling, storage, networking, RBAC, certificates and more), it answers in four parts: Cause, Diagnose (commands), Fix (commands or YAML), Verify.

A fresh LoRA with v1's configuration (r=16 / Ξ±=32), trained on the same data as the 0.5B v2. v1's adapter only fits the 0.5B base, so this model does not continue it.

Prototype β€” verify before you run anything. This model reliably produces the right shape of answer, and it gets some common failures right. It also states wrong root causes and wrong fixes with full confidence (see Known failures below). Treat every command and YAML as a suggestion to review, never as something to apply to a cluster unchecked.

Training

  • Data: 1270 train / 69 held-out chat examples. 2 epochs, learning rate 0.0001, LoRA r=16 / Ξ±=32 on all attention and MLP projections. Loss is on assistant tokens only; bf16 on an NVIDIA GB10 (DGX Spark).
  • How the examples were made: Kubernetes troubleshooting sources covering 47 error classes were collected and chunked, and a local teacher model (mostly gpt-oss:20b, some Qwen3-8B) wrote 1–3 grounded examples per chunk. Every example was checked (real kubectl subcommands, parseable YAML, no near-duplicates) and each error class was capped at 5% of the set.
Source License Use
kubernetes.io debug and task docs (English) CC BY 4.0 grounding for generated examples
Stack Overflow / Server Fault Q&A CC BY-SA 4.0 grounding for generated examples
Scoutflo SRE Playbooks (K8s) MIT grounding for generated examples
kubernetes/* GitHub issue threads GitHub ToS grounding only, paraphrased
Vendor troubleshooting guides proprietary grounding only, paraphrased β€” no text copied
jalpan04/devops-sft-dataset (Kubernetes questions only) Apache-2.0 direct rows
spacezenmasterr/k8s-sft-cmd-en MIT direct rows

v1's AnveshGummala/k8s-troubleshooting-customdsv3 is not used: its license is undeclared. Research-paper Q&A rows that v1's source mixed into its "Kubernetes" rows were removed.

Evaluation

Two measures, run on 2026-09-25 with the same code for every model:

  • Rubric ↑: over 40 canonical scenarios, the share of required root-cause and fix keywords each answer contains. It is coarse. It rewards naming the right things, not being right, so read it next to Known failures.
  • Loss ↓: loss on held-out assistant tokens, both on new-style examples and on v1-style rows.
model rubric ↑ loss, new data ↓ loss, v1-style ↓
Qwen2.5-0.5B-Instruct (no adapter) 0.487 1.581 2.082
t4tarzan/DKube-instruct (v1, 0.5B) 0.367 1.689 1.664
DKube-instruct-v2-0.5B 0.562 1.173 1.593
Qwen2.5-1.5B-Instruct (no adapter) 0.608 1.332 2.072
DKube-instruct-v2-1.5B 0.642 0.944 1.610
Qwen2.5-7B-Instruct (no adapter) 0.825 1.435 3.069
DKube-instruct-v2-7B (this model) 0.821 0.757 1.393

Known failures

Read from this model's own answers to the canonical scenarios:

  • Exit codes 126 and 139: calls both "SIGKILL". 126 means the command was found but is not executable (check the entrypoint's permissions); 139 is SIGSEGV, a segfault in the process.
  • Pod stuck Terminating: blames resource constraints; the usual cause is a finalizer, or a node that is gone (kubectl get pod -o jsonpath='{.metadata.finalizers}', then a reviewed --grace-period=0 --force).
  • RBAC Forbidden: the diagnosis is right, but the fix it proposes is a ClusterRoleBinding that reads all resources; least privilege is a namespaced Role + RoleBinding for pods with get/list/watch.
  • On the rubric it ties the untrained Qwen2.5-7B-Instruct (0.821 vs 0.825, 5 wins / 6 losses / 29 ties). At this size the fine-tune mostly teaches the answer format and the house style (held-out loss 1.435 β†’ 0.757), not new facts.
  • Scored well (28 of 40 perfect): exit 137 β†’ OOMKilled, 401 β†’ imagePullSecrets, wrong tag, rate limits, exec format error, CreateContainerConfigError, init containers, every scheduling case, PVC / FailedMount / Multi-Attach, NotReady, DiskPressure, probes, ProgressDeadlineExceeded, Service endpoints, DNS, ingress 502, x509, webhooks, quota, stuck namespaces, localhost:8080 refused.

Usage

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "Qwen/Qwen2.5-7B-Instruct"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto")
model = PeftModel.from_pretrained(model, "dkubeio/DKube-instruct-v2-7B")

msgs = [
    {"role": "system", "content": "You are DKube-instruct, a Kubernetes troubleshooting assistant. Diagnose the root cause, give exact kubectl commands or YAML fixes, and say how to verify."},
    {"role": "user", "content": "Pod is CrashLoopBackOff with exit code 137. How do I fix it?"},
]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt")
print(tok.decode(model.generate(ids, max_new_tokens=400)[0][ids.shape[1]:], skip_special_tokens=True))

The system prompt above is the one the model was trained with; answers are best with it.

License

The adapter is released under Apache-2.0. Use of the base model follows its own license (Qwen/Qwen2.5-7B-Instruct: Apache-2.0).

Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for dkubeio/DKube-instruct-v2-7B

Base model

Qwen/Qwen2.5-7B
Adapter
(2813)
this model