Instructions to use nima1/bashcraft-qwen2.5-coder-1.5b-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use nima1/bashcraft-qwen2.5-coder-1.5b-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Coder-1.5B-Instruct") model = PeftModel.from_pretrained(base_model, "nima1/bashcraft-qwen2.5-coder-1.5b-lora") - Notebooks
- Google Colab
- Kaggle
Bashcraft LoRA adapter
Educational suggestion model. It does not execute commands and is not a safe shell agent.
Bashcraft adapts a small instruction model to return a Bash command suggestion or a question for missing information from an English request and explicit context. The supplied CLI validates the JSON reply and does not execute suggestions. A well-formed or syntax-valid command can still select the wrong files, change unrelated state, or fail the requested task.
This repository distributes LoRA weights, not the base model. The required base
is Qwen/Qwen2.5-Coder-1.5B-Instruct
at exact revision 2e1fd397ee46e1388853d2af2c993145b0f1098a, with its matching
tokenizer. The base weights remain frozen; the adapter adds trained attention
projection matrices.
Release identity and selection
Selected adapter: m4-train-v1, seed 42, one epoch on all 2,000 frozen
training records. Its original training-manifest digest is
b6bd082ea1b000d3aa53ccd2a74535fe29d98098c820f95c140ee3ca578be31c.
The portable manifest has a new digest because machine-specific metadata is
removed; inference.json binds those released bytes. The prompt contains three
exact training examples selected before test scoring.
Validation selection: three examples scored 57/100 executable tasks on the base model versus 53/100 with six. The 2,000-example adapter scored 72/100 with instructions alone versus 55/100 for 500 examples. Adding the selected three examples scored 74/100, so the preregistered success-first rule selected it for release. Both 2,000-example prompt variants had one unintended change. Ambiguity behavior regressed: the selected prompt returned no clarification responses on ten ambiguous validation requests, versus nine with instructions alone. This choice is not a claim of uniformly better behavior.
Prepared destinations are nima1/bashcraft-qwen2.5-coder-1.5b-lora and
nima1/bashcraft-data. The project release record records immutable Hub commit
IDs after publication. For reproduction, select that exact revision when
downloading; do not rely on a later moving main.
The artifact files include:
adapter_model.safetensorsandadapter_config.json: verified LoRA state and pinned base identity.manifest.json: portable adapter identity and original evidence digests.inference.json: the selected prompt, greedy decoding settings, and adapter hash.training-selection.json: exact selected training IDs/count/hash, parent corpus hash, training recipe, and source/dependency identities.release-manifest.json: hashes of distributed files.uv.lock,LICENSE,NOTICE,licenses/Qwen.LICENSE, andATTRIBUTION.md: dependency identity, license text, and attribution.
The software artifacts are bashcraft-0.1.0-py3-none-any.whl and
requirements-ml.txt. The latter exports the exact ML dependency lock with
package hashes and the PyPI/PyTorch CUDA 13.0 indexes. A fresh environment
installed all 80 dependencies from cache and the wheel successfully; the CLI,
CPU diagnostic and package checks passed. This is a cached-dependency install,
not an offline bundle of the large dependencies or base model.
After downloading an immutable revision into bashcraft-model, install on
Linux with a supported NVIDIA CUDA/BF16 device:
uv venv --python 3.13.16 bashcraft-env
uv pip sync --require-hashes --python bashcraft-env/bin/python \
bashcraft-model/requirements-ml.txt
uv pip install --no-deps --python bashcraft-env/bin/python \
bashcraft-model/bashcraft-0.1.0-py3-none-any.whl
cd bashcraft-model
../bashcraft-env/bin/bashcraft suggest --config inference.json \
--request 'Print only the final two lines of activity.txt without changing it.'
The first inference may download the pinned base/tokenizer. The suggestion interface requires at least 4,096 MiB reported free before loading; actual peak use depends on prompt length and other workloads. Do not equate that preflight threshold with a universal hardware guarantee.
Training data and procedure
The frozen parent training corpus has 2,000 original agent-assisted examples:
1,900 command tasks and 100 clarification requests, across 20 command and 10
clarification families. It uses deterministic template/parameter expansion,
not an imported translation corpus or a teacher-model sampling run. The smaller
candidate is an exact nested 500-record subset with 475 commands and 25 clarifiers.
The published dataset contains the full parent; training-selection.json identifies
which records trained this adapter. Validation and test targets are absent from
the training-data release.
All 1,900 parent training references passed isolated syntax/outcome checks; this validates reference tasks, not model behavior. An independent AI reviewer inspected 100 training descriptions and all clarification targets. Synthetic templates, correlated family variants, limited primitives, and small disposable fixtures restrict generalization claims.
The preregistered adaptation recipe is BF16 backbone plus PEFT LoRA on
q_proj, k_proj, v_proj, and o_proj: rank 16, alpha 32, dropout 0.05.
Only 4,358,144 adapter parameters are trainable. One epoch uses maximum sequence
length 512 without truncation, microbatch 2 and accumulation 8, constant learning
rate 1e-4, AdamW weight decay 0.01, betas (0.9, 0.999), epsilon 1e-8, and
gradient clipping at norm 1. Gradient checkpointing, quantization, and CPU offload
are disabled. Loss covers assistant JSON and its ending token, masks the prompt
and padding, and weights accumulation by supervised token count.
The selected run used 125 optimizer updates, 539,346 input tokens and 71,320
supervised assistant tokens. Measured training time was 41.026757 seconds;
total run time 45.438946 seconds; checkpoint time 0.456584 seconds. Peak PyTorch
allocated/reserved memory was 6,534,043,136 / 8,629,780,480 bytes on an RTX5090
with other workloads left running. These timings do not include data engineering
or evaluation. The initial microbatch-four pilot failed with CUDA OOM before an
optimizer update; microbatch two with accumulation eight succeeded. A fresh full
run then used the successful recipe. The main training code revision was
11bae45 with runtime source hashes preserved in training metadata.
Evaluation
On the frozen authored final test, the selected adapter solved 143/200 (71.5%) executable tasks versus 95/200 (47.5%) for the base model with the identical three-example prompt. The paired gain is 24.0 percentage points, with a 95% task-group bootstrap interval of 13.1–35.4 points (25 groups). This supports a gain on this benchmark, not general shell reliability.
| Method | Task success | Format valid | Checked syntax | Unexpected changes / observable states |
|---|---|---|---|---|
| Base, instructions | 1/200 | 12/220 | 3/3 | 0/3 |
| Base, three examples | 95/200 | 214/220 | 193/194 | 7/193 |
| Adapter seed42, instructions | 96/200 | 220/220 | 194/198 | 3/194 |
| Adapter seed42, three examples (release) | 143/200 | 218/220 | 187/198 | 1/187 |
| Adapter seed43, three examples | 141/200 | 219/220 | 188/199 | 1/188 |
| Adapter seed44, three examples | 139/200 | 220/220 | 191/200 | 2/191 |
The three fixed training seeds scored 71.5%, 70.5%, and 69.5%; their mean is 70.5%. Seed42 remains the release regardless of these test scores. The primary selected adapter solved 129/150 familiar tasks (86.0%), 10/30 held-out compositions (33.3%), and 4/20 compositions with a new option (20.0%). The much weaker held-out results limit generalization claims.
Selected-model format validity was 218/220 and checked syntax 187/198. One observed case created an unintended file after dropping literal apostrophes from a filename. Other failures include missing depth limits, regex substitution where literal replacement was requested, and incorrect pipelines. Thirteen of the 200 executable cases lack a complete observable-state grade; they are not evidence of no side effects. The 20 clarification cases were intentionally unexecuted. Of 20 ambiguous requests, the model returned 16 commands and four questions. A disclosed author-assisted, non-blinded review rated three questions adequate and one misleading. None received automatic success credit.
Generation-only median/p95 was 0.2240/0.2986 seconds on this RTX5090, including the first cold call but excluding loading and isolated grading. Peak PyTorch inference allocation was 3,213,895,680 bytes. These timings and memory are local measurements, not deployment guarantees.
The six-method protocol was frozen and committed as 023dfd3 before final
inference. All first answers were retained. The comparison independently
regraded saved outcomes and verified config, data, execution, source and adapter
hashes. The 500/2,000 comparison is at constant epochs, so it changes compute as
well as data size. The optional 5,000 run was not performed: it would repeat the
same generator families; this is not evidence that additional examples cannot help.
The primary preregistered comparison is seed-42 adapter versus base under the
same validation-selected few-shot prompt. Its paired 95% bootstrap interval uses
10,000 draws, seed 20261004, resampling split_group; clarification cases are
excluded from executable success. Seeds 43 and 44 assess variation without selecting
the best test seed. Report a gain only if the primary interval excludes zero in
the positive direction, and retain null or negative results.
Task success requires output/state assertions and no unpermitted changes. Format and syntax denominators differ and must remain explicit. Clarification quality requires separate qualitative review. Generation latency excludes model loading and sandbox grading, and includes the first cold call. PyTorch memory allocation does not equal total GPU occupancy.
Suggestion-only use
The installation recipe above passed a fresh cached-dependency CPU check using the local release wheel. The installed wheel also loaded the local portable adapter offline and produced a first file suggestion that passed its isolated fixture. This card was prepared before publication: downloaded-artifact verification had not yet run. The project release report records that later check against exact Hub revisions when completed. The following is the suggestion interface; its output is not predetermined. A file-hash check alone does not prove successful inference.
With Bashcraft and its locked ML dependencies installed, run from the downloaded
model directory containing inference.json and manifest.json:
../bashcraft-env/bin/bashcraft suggest --config inference.json \
--request 'List regular .log files directly in the current directory.'
The portable config's adapter.path is . and resolves against the working
directory. From another directory, use the absolute-path
downloaded-inference.json produced by Bashcraft's verified-download workflow,
or construct an equivalent local config while preserving the manifest hash.
Do not assume --config /some/path/inference.json also changes the working
directory. The runtime requires a supported NVIDIA CUDA device with BF16 support;
there is no automatic CPU fallback.
A valid reply has one of these shapes:
{"status":"command","command":"..."}
{"status":"clarify","question":"..."}
These shapes are illustrations, not promised outputs for the example request.
Malformed responses fail without repair or execution. Optional --context accepts
a JSON task context, and --metrics includes model identity and generation metrics.
No output should be piped automatically into a shell.
Scope, limitations, and attribution
Intended use is education, controlled experimentation, and reviewed command suggestions for the bounded Bash/GNU/local-Git task scope. Docker/npm tasks, network operations, production repositories, autonomous execution, and security assurance are outside the evidence. The model can invent paths, quote arguments incorrectly, omit selection boundaries, ask unnecessary questions, or misunderstand state changes. Human inspection and an appropriate disposable test environment remain necessary before a suggestion is used.
The project repository is kkarimi/bashcraft. It is currently private and requires access. Compact reports, tutorials, and experimental records are kept there; local raw snapshots are not distributed by this model repository. This model card includes the measured comparison and installation procedure; the distributed wheel and dependency export permit inference without access to that private source repository. The detailed study and tutorials require repository access.
Original Bashcraft adapter material is released under Apache-2.0. The base model
is copyright 2024 Alibaba Cloud and is distributed under Apache-2.0; its license
is preserved in licenses/Qwen.LICENSE. See LICENSE, NOTICE, and
ATTRIBUTION.md. Base-model weights are downloaded separately and are not part
of this adapter repository.
- Downloads last month
- 14
Model tree for nima1/bashcraft-qwen2.5-coder-1.5b-lora
Base model
Qwen/Qwen2.5-1.5B