EasyCommand 1.5B: trainable checkpoint

The selected A3 EasyCommand checkpoint, supplied as a complete merged BF16 Hugging Face model and its original cumulative FP32 LoRA adapter. It generates GNU/Linux Bash commands from English requests, using COMMAND JSON.

This is a fine-tuned model derived from Qwen/Qwen2.5-Coder-1.5B-Instruct, pinned to revision 2e1fd397ee46e1388853d2af2c993145b0f1098a. It is not the untouched upstream model. The merged weight file is 3,087,467,144 bytes (3.09 GB).

Files and uses

  • model.safetensors, configuration and tokenizer: standalone Transformers inference or a starting point for a new fine-tuning run.
  • adapter/: the original trained LoRA weights and portable PEFT configuration, for continuing the existing adapter with its pinned base.
  • system-prompt.txt: the exact evaluated serving prompt.
  • training.json: training recipe and initializer lineage.
  • manifest.json and SHA256SUMS: file identity and integrity.

The matching GGUF release provides the evaluated quantized deployment files. The dataset is available separately under Apache-2.0.

Load the merged model

The release was checked with Transformers 4.57.6, PyTorch 2.9.1 and PEFT 0.21.1. Install PyTorch for your hardware, plus transformers, accelerate, safetensors and huggingface_hub; add peft when using the adapter.

import json
from pathlib import Path
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "dirac-run/ec-1.5b"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo, dtype=torch.bfloat16, device_map="auto"
)
model.eval()
system = Path(hf_hub_download(repo, "system-prompt.txt")).read_text()
messages = [
    {"role": "system", "content": system},
    {"role": "user", "content": "print the system uptime"},
]
inputs = tokenizer.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    enable_thinking=False, return_dict=True, return_tensors="pt"
).to(model.device)
with torch.inference_mode():
    output = model.generate(
        **inputs, do_sample=False, max_new_tokens=256,
        eos_token_id=151645, pad_token_id=151643
    )
response = tokenizer.decode(
    output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True
)
print(json.loads(response)["value"])

This prints a proposed command; it does not execute it. Use Qwen2.5 ChatML without a Qwen3 thinking prefix. Keep the system prompt, chat template and greedy decoding when reproducing the documented deployment profile. JSON is learned behavior; no grammar is applied. Other backends, numeric formats or kernels can change outputs.

Continue the original LoRA adapter

import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM

base = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen2.5-Coder-1.5B-Instruct",
    revision="2e1fd397ee46e1388853d2af2c993145b0f1098a",
    dtype=torch.bfloat16,
)
model = PeftModel.from_pretrained(
    base, "dirac-run/ec-1.5b", subfolder="adapter", is_trainable=True
)

Use the pinned original upstream model with the existing adapter. Applying that adapter to the merged EasyCommand model would apply the learned change twice. To add a new adapter to the merged model instead, load the root checkpoint and create a fresh LoRA configuration. Both approaches start a new optimization run: optimizer, scheduler and RNG states from the old run are not included.

Training and conversion

One weighted epoch of 474,635 presentations took 14,833 LoRA updates from the upstream weights. The selected adapter then received 200 repair/replay updates from that trained state; its saved weights include both stages.

Both stages used rank 32, alpha 64, dropout 0.05 and all q/k/v/o attention and gate/up/down MLP projections, with effective batch 32, a frozen BF16 base and FP32 adapters. AdamW used warmup and cosine decay. The parent peak LR was 0.0001; the continuation used 5e-6, ten warmup updates and 50% repair / 50% replay sampling from 2,820 source rows. The parent ran on H100, the continuation on A40. Assistant answer and EOS tokens were supervised; prompt/padding tokens were masked.

The longer parent training prompt differs from the 177-byte serving/continuation prompt. The released dataset is a deduplicated union of 401,975 pairs, including other repair experiments. One flat pass over it does not reproduce historical repetition, sampling or this model's exact exposure.

The merged checkpoint computes base plus scaled LoRA delta in FP32, then rounds once to BF16. Its pre-rounding merge was verified byte-for-byte against the FP32 merge used to produce the selected GGUF. Tied embedding/output weights are preserved. The adapter binary remains the original FP32 file; only its non-portable base path was replaced with the pinned HF base ID/revision.

Evaluation and limitations

The per-quant ALFA-updated and six-panel internal results are reported on the matching GGUF card. The merged BF16 export has passed loading, generation and tensor checks. Adapter loading and finite, nonzero training gradients were also checked without an optimizer update. This newly rounded BF16 export has not received a fresh full benchmark run; do not transfer a GGUF score to it.

ALFA-updated is a separately documented benchmark variant, not original ALFA. ALFA failures and internal panels informed data repair and checkpoint selection. Those results are development measurements, not an untouched independent test.

The model targets English requests, Bash and GNU/Linux utilities. It does not inspect the live filesystem or know which tools are installed. It gives a best-effort command rather than clarification or inability responses. Commands can be incorrect, incomplete or destructive; inspect them before execution. The ec application is published separately; the Transformers example here also works independently of it.

Published original-ALFA results (different protocol)

These externally reported results provide context. They are not directly comparable with our ALFA-updated measurements. Sources checked on 5 October 2026.

Model/configuration Reported download size Original-ALFA pass rate Source
GPT-4o, cloud API (published reference) — 73.0% whatisit benchmarks
nl2sh-3b Q4_K_M 1.9 GB 65.7% whatisit benchmarks
Community nl2sh-qwen25-coder-1.5b Q4_K_M 941 MB 65.67% Community model card
whatisit / nl2sh-1.5b Q4_K_M 941 MB 62.0% whatisit benchmarks
Qwen2.5-Coder-7B, untuned 4.4 GB 61.3% whatisit benchmarks
Qwen2.5-Coder-1.5B, untuned 941 MB 54.0% whatisit benchmarks

The local-model sources report 300 tasks, temperature 0, a 64-token output cap, and the unmodified upstream scorer with embedding threshold 0.75. GPT-4o's 73.0% is the benchmark paper's published reference, not a run by us or a measurement under that local serving profile. Sizes retain the authors' units and rounding; the untuned rows' quantization is not specified in the source table.

The community author also remeasured whatisit at 59.0% on their own rig, versus its upstream published 62.0%. Both are external original-ALFA measurements, distinct from our 191/300 (63.67%) ALFA-updated result. No EC-versus-external ranking across these two grading protocols is implied.

License and integrity

The model and this documentation are licensed under Apache-2.0, with upstream/project attribution in NOTICE. The adapter has the same terms. Verify downloaded files from the repository root with sha256sum --check SHA256SUMS.

Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dirac-run/ec-1.5b

Finetuned
(216)
this model

Dataset used to train dirac-run/ec-1.5b