File size: 8,866 Bytes
bc6c786
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
---
base_model: LiquidAI/LFM2-700M
library_name: transformers
pipeline_tag: text-generation
license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2-700M/blob/main/LICENSE
language:
- en
tags:
- lfm2
- linux
- bash
- shell
- command-generation
- text-generation
datasets:
- jiacheng-ye/nl2bash
- mecha-org/linux-command-dataset
---

# thealper2/lfm2-700m-linux-command

LiquidAI/LFM2-700M fine-tuned to map a natural-language Linux task to a single shell
command. The target output is the command only; no explanation is produced.

## Model details

| Field | Value |
| --- | --- |
| Base model | LiquidAI/LFM2-700M |
| Architecture | `Lfm2ForCausalLM` β€” hybrid, 16 layers: full attention at `[2, 5, 8, 10, 12, 14]`, gated short convolution elsewhere |
| Parameters | 742,489,344 total (641,826,048 non-embedding) |
| Hidden size / heads / KV heads | 1536 / 24 / 8 |
| Vocabulary | 65,536 |
| Context length (base) | 128,000 |
| Training precision | bfloat16 |
| Fine-tuning method | full |
| Task | NL instruction β†’ shell command |

## Prompt format

The model uses the LFM2 ChatML-style chat template and was trained with no
system prompt. Apply the template rather than constructing the string manually.

```
<|startoftext|><|im_start|>user
Find which process is using port 8080.<|im_end|>
<|im_start|>assistant
lsof -i :8080<|im_end|>
```

Two LFM2 tokenizer details matter:

- The chat template emits `bos_token` itself and `add_bos_token` is `true` in
  `tokenizer_config.json`. Tokenise templated text with
  `add_special_tokens=False`, or the sequence gets a duplicated BOS.
- Decode with `clean_up_tokenization_spaces=False`; the BPE cleanup step strips
  spaces around punctuation and can corrupt shell commands.

Special tokens: BOS `<|startoftext|>` (1), EOS `<|im_end|>` (7), PAD `<|pad|>` (0).

## Usage

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "thealper2/lfm2-700m-linux-command"
tokenizer = AutoTokenizer.from_pretrained(model_id, clean_up_tokenization_spaces=False)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16).to("cuda").eval()

messages = [{"role": "user", "content": "Find which process is using port 8080."}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)

output = model.generate(
    **inputs,
    max_new_tokens=128,
    do_sample=False,                      # deterministic decoding
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.pad_token_id,
)
command = tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip()
print(command)  # lsof -i :8080
```

Greedy decoding (`do_sample=False`) is the intended configuration: the task has a
single intended answer and sampling only adds variance.

## Training data

| Source | Raw rows | Schema |
| --- | --- | --- |
| [`jiacheng-ye/nl2bash`](https://huggingface.co/datasets/jiacheng-ye/nl2bash) | 9,305 | `nl`, `bash` |
| [`mecha-org/linux-command-dataset`](https://huggingface.co/datasets/mecha-org/linux-command-dataset) | 8,669 | `input`, `output` |

Both were normalised to `{instruction, command, source}` and then:

1. **Cleaned** β€” markdown fences and `$`/`#` prompt prefixes stripped, whitespace
   runs collapsed outside quoted strings, records with unbalanced quotes, prose
   instead of a command, or no parseable utility dropped. Commands themselves
   were never rewritten.
2. **Deduplicated** β€” exact `(instruction, command)` duplicates removed. Rows
   sharing a command with a different instruction, or an instruction with a
   different valid command, were kept deliberately.
3. **Balanced** β€” `find` was 35.2% of the raw corpus. It was capped per-utility
   using a diversity-aware ordering that retains every distinct flag signature
   before retaining any repeat, lowering its share to ~19%.
4. **Split** β€” 90/5/5, grouped by instruction *template* (filenames, paths,
   numbers and quoted literals abstracted) so templated paraphrases cannot
   straddle the train/test boundary.

Split sizes: 11756 train / 652 validation / 653 test.
Leakage checks report zero overlap across splits at the exact-pair,
instruction and instruction-template level.

## Training configuration

| Hyper-parameter | Value |
| --- | --- |
| Method | full |
| Epochs | 3 |
| Learning rate | 3e-05 |
| Scheduler | cosine |
| Warmup ratio | 0.03 |
| Weight decay | 0.01 |
| Optimizer | adamw_torch_fused |
| Per-device batch size | 16 |
| Gradient accumulation | 2 |
| Max sequence length | 128 |
| Precision | bfloat16 |
| Seed | 42 |

Loss is computed on the assistant turn only; prompt tokens are masked with
`-100`.

`max_length` was chosen from the tokenised length distribution of the corpus
(mean 37.7, median 34, p90 57, p95 66, p99 87, max 403) β€” a 128-token budget
covers 99.94% of examples.

### Run record

| Field | Value |
| --- | --- |
| Final training loss | 0.516 |
| Validation loss | 0.6445 |
| Training time | 425.0 s |
| Peak VRAM | 8.95 GB |
| GPU | NVIDIA GeForce RTX 5060 Ti |
| torch / transformers | 2.11.0+cu128 / 5.17.0 |

## Evaluation

Measured on the held-out test set with greedy decoding.

| Metric | Base LFM2-700M | Fine-tuned | Delta |
| --- | --- | --- | --- |
| Exact match | 0.0061 | 0.2910 | +0.2849 |
| Normalised exact match | 0.0061 | 0.2910 | +0.2849 |
| Structural match | 0.0107 | 0.3032 | +0.2925 |
| Command validity | 0.4763 | 0.9939 | +0.5176 |
| Token F1 | 0.1195 | 0.6470 | +0.5275 |
| Primary-utility accuracy | 0.1807 | 0.8377 | +0.6570 |
| Prose-output rate | 0.3783 | 0.0000 | -0.3783 |

Metric definitions:

- **Exact match** β€” string equality after stripping surrounding whitespace.
- **Normalised exact match** β€” equality after collapsing whitespace runs outside
  quotes and removing a trailing `;`.
- **Structural match** β€” utility, flag multiset (short-flag bundles expanded for
  utilities that use them) and operand sequence compared per pipeline segment.
  Recognises `ls -la` == `ls -al`.
- **Command validity** β€” the output parses under a bash grammar parser
  (`bashlex`), not membership in a list of known utilities.
- **Token F1** β€” token-level overlap, as partial credit.
- **Primary-utility accuracy** β€” the first utility matches the reference.
- **Prose-output rate** β€” fraction of outputs that read as an explanation rather
  than a command.

## Limitations

- **Structural match is not a semantic oracle.** It compares command *shape*. It
  cannot tell that `find . -name '*.py'` and a shell glob achieve the same
  result, and it does not reason about flag semantics. It is an upper bound on
  exact match, not semantic accuracy.
- **Exact match understates correctness.** Many Linux tasks have several valid
  answers; the test set carries one reference each.
- **Source-distribution bias.** `nl2bash` is `find`-heavy and composition-heavy;
  `mecha-org/linux-command-dataset` is templated and single-utility-heavy.
  Per-source metrics differ and are reported separately in the project reports.
- **Distribution shift.** Commands reference paths, hosts and variables that
  appear in the training corpora (`/path/to/...`, `$source`). Outputs may embed
  those placeholders instead of the user's real paths.
- **Short outputs only.** Trained at a 128-token budget; long multi-stage scripts
  are out of distribution.
- **No verification of correctness or safety at generation time.** The model can
  produce syntactically valid but wrong β€” or destructive β€” commands.

## Intended use and safety

Intended for generating candidate shell commands for review, and as the command
generator of a sandboxed terminal agent.

**Do not execute generated commands directly on a host.** The project this model
comes from executes commands only inside a disposable Docker container started
with `--network none`, `--read-only`, `--cap-drop ALL`,
`--security-opt no-new-privileges`, a non-root user, no host bind mounts,
bounded CPU/memory/PIDs and a wall-clock timeout, and screens commands against a
destructive-pattern list and a read-only allowlist before running them.

## License

Inherits the LFM Open License v1.0 of the base model, LiquidAI/LFM2-700M. Dataset
licenses apply to the training data: `mecha-org/linux-command-dataset` is
Apache-2.0; `nl2bash` derives from the
[TellinaTool/nl2bash](https://github.com/TellinaTool/nl2bash) corpus.

## Citation

The NL2Bash corpus:

```bibtex
@inproceedings{LinWZE2018:NL2Bash,
  author    = {Xi Victoria Lin and Chenglong Wang and Luke Zettlemoyer and Michael D. Ernst},
  title     = {NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System},
  booktitle = {LREC 2018},
  year      = {2018}
}
```