Instructions to use Irfanuruchi/Qwen3-14B-BuildEng with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Irfanuruchi/Qwen3-14B-BuildEng with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-14B") model = PeftModel.from_pretrained(base_model, "Irfanuruchi/Qwen3-14B-BuildEng") - Notebooks
- Google Colab
- Kaggle
Qwen3-14B-BuildEng
Qwen3-14B-BuildEng is my Building Engineering fine-tune of Qwen/Qwen3-14B. The goal was not to turn a 14B model into a replacement for an engineer. I wanted a smaller model that is actually useful for the kind of compact engineering work that comes up all the time: HVAC and ventilation, building physics, duct and pipe calculations, energy and quantity calculations, electrical building-services questions, unit handling, basic structural calculations, sanity checks, and spotting missing or contradictory inputs.
This repository contains the selected u096 LoRA adapter. The canonical evaluated model is NF4(Qwen3-14B) + u096 LoRA. That distinction matters because the merged BF16 and GGUF builds are deployment branches, not byte-for-byte or behaviorally identical representations of the model used for the final held-out evaluation.
Model setup
Base model: Qwen/Qwen3-14B
Task: causal language modeling
Adapter: LoRA
Rank: 16
Alpha: 32
Dropout: 0.05
Target modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
The model was trained with QLoRA using an NF4 4-bit base, BF16 compute, microbatch 2, gradient accumulation 8, an effective batch of 16 packed samples per optimizer update, native torch.optim.AdamW, base learning rate 5e-5, 16 warmup updates followed by a constant learning rate, gradient clipping at 1.0, full reentrant gradient checkpointing, dynamic sequence trimming, and an assistant-only weighted target-token objective. The training seed was 20261001.
Checkpoint u096 was selected from the frozen validation and qualitative evaluation process. Later checkpoints kept pushing validation loss lower, but several frozen out-of-sample engineering behaviors got worse. I kept u096 instead of selecting a checkpoint only because its validation number looked better.
Final held-out result
The closed held-out test contained 1007 rows from 619 unique user groups. The base model scored NLL 3.217899 with perplexity 24.9756; the selected u096 adapter scored NLL 0.135998 with perplexity 1.14568. On that specific assistant-target token metric, this is a 95.77% relative NLL reduction.
That number should be read exactly for what it is: a result on the frozen held-out objective used for this project, not a claim that the model is “95.77% better” at engineering in general. The held-out test was closed after the final run and was not reused to choose deployment quantizations.
The preserved records are in research/BUILDENG-CHECKPOINT-SELECTION-001.json, research/BUILDENG-FINAL-MODEL-RECORD-001.json, research/BUILDENG-FINAL-MODEL-RECORD-001-SHA256SUMS.txt, and research/BUILDENG-FINAL-TEST-U096-003.json.
What it is meant for
This model works best on focused Building Engineering questions where the inputs are reasonably clear and the answer can be checked. Typical examples are ACH and airflow calculations, duct or pipe velocity, thermal resistance and U-values, fan or heat-pump power, building-services electrical arithmetic, water and energy calculations, basic statics, unit conversion, and engineering sanity checks.
It was also trained to be conservative when information is missing or contradictory instead of quietly inventing a standard, clause, design value, or assumption. That behavior is important to me because a model that confidently fills in missing engineering data is more dangerous than one that simply says the problem is underspecified.
This is still a 14B model. It can make arithmetic mistakes, lose track of a reference in a longer conversation, drift on some diagnostic questions, or start hallucinating formulas when pushed into a large multi-stage design problem. Complex code-based structural design, regulatory compliance, safety-critical design, and final professional decisions still need proper standards, project data, and engineering review.
System prompt used in evaluation
You are a Building Engineering assistant. Answer accurately and concisely. Show calculations and units when relevant. Do not invent missing engineering inputs, standards clauses, or code requirements. If the information is insufficient or contradictory, state that clearly.
Loading with Transformers + PEFT
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
base_id = "Qwen/Qwen3-14B"
adapter_id = "Irfanuruchi/Qwen3-14B-BuildEng"
quant_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
tokenizer = AutoTokenizer.from_pretrained(base_id)
base = AutoModelForCausalLM.from_pretrained(
base_id,
quantization_config=quant_config,
device_map="auto",
)
model = PeftModel.from_pretrained(base, adapter_id)
messages = [
{
"role": "system",
"content": (
"You are a Building Engineering assistant. Answer accurately and concisely. "
"Show calculations and units when relevant. Do not invent missing engineering "
"inputs, standards clauses, or code requirements. If the information is "
"insufficient or contradictory, state that clearly."
),
},
{
"role": "user",
"content": "A room is 8 m by 6 m by 3 m and requires 2.5 ACH. What airflow is required?",
},
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=512,
do_sample=False,
)
answer = tokenizer.decode(
output[0][inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
)
print(answer)
GGUF builds
The GGUF deployment builds are published separately at Irfanuruchi/Qwen3-14B-BuildEng-GGUF.
All three deployment variants were run through the same frozen 24-case qualitative suite. Q8_0 finished at 20 PASS / 1 PARTIAL / 3 FAIL and is the recommended fidelity-oriented GGUF. Q5_K_M finished at 19 / 1 / 4, while Q4_K_M also finished at 19 / 1 / 4 and is the most compact option of the three.
Those results are deliberately kept separate from the canonical LoRA held-out result. The GGUF files come from the merged deployment branch, so I do not claim that they have the same behavioral identity or the same held-out metrics as NF4(base) + u096 LoRA.
One interesting thing from the quantization work is that the behavior was not simply “fewer bits = every case gets worse.” Q8, Q5 and Q4 sometimes took different generation paths, and Q4 could recover one part of a calculation while making a different mistake later in the same answer. That is why Q8_0 is the default recommendation when memory allows it, while Q5_K_M and Q4_K_M are there for smaller deployments.
Provenance
The public adapter weights are byte-identical to the selected frozen u096 research adapter.
adapter_model.safetensors SHA256:
15a367fb0b8c0b72c9ad0daad1fa532ca8dba4c0dc686c10578832fd95e8223a
The only publication-time change to the adapter metadata was normalizing base_model_name_or_path from the local training path /home/eagle/AI-Models/huggingface/Qwen3-14B to the public identifier Qwen/Qwen3-14B. No model weights were modified. That transformation is recorded in PUBLICATION-METADATA-TRANSFORM-001.txt.
The release keeps the selected adapter, merged BF16 branch, BF16 GGUF master, Q8_0, Q5_K_M, Q4_K_M, evaluation records, quantization records, and SHA256 manifests as separate artifacts. The final held-out test was not rerun for deployment-format selection. Research validation and deployment characterization were treated as separate stages on purpose.
Disclaimer
This model is for research, education, and engineering-assistant use. It can be useful, but it can also be wrong. Do not use it alone for safety-critical engineering decisions, structural approval, regulatory compliance, construction authorization, contractual design decisions, or final professional calculations. Important outputs should be checked against the relevant standards, project information, calculation method, and qualified engineering judgment.
- Downloads last month
- 22