ModernBERT-bash-classifier

This model is a fine-tuned version of answerdotai/ModernBERT-base on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 0.1192
  • F1: 0.9615
  • Precision: 0.9615
  • Recall: 0.9615
  • Accuracy: 0.9893

Model description

Bash command classification model, for classifying a bash command as either safe or unsafe based on what it does.

We classify a bash command as unsafe based on the following criteria:

  • Downloads data from the internet[^1]
  • Sends data from the machine over the network[^1]
  • Permanently modifies files outside the current project directory[^2]
  • Deletes files outside the current project directory[^2]
  • Modifies cloud resources (via CLIs like az or tf)
  • Creates persistent, resource consuming processes on the machine like cron jobs

[^1]: Git commands are an exception to this rule.

[^2]: We treat files within the current project directory as safe, since an agent working in a project is expected to mutate its contents.

Intended uses & limitations

Created to be used as a auto-mode classifier for coding agent harnesses, so users only manually approve bash commands flagged as unsafe.

The model may not generalize well to bash commands incorporating unknown CLIs or commands that use raw code blocks like Python within heredocs.

Training and evaluation data

Fine-tuned on bash tool calls collected from Claude Code and Codex traces. Any PII in the commands such as directory names were anonymised. The total dataset consists of 3,738 bash calls.

Training procedure

Fine-tuned on bash commands formatted as CWD: [CWD]\nCOMMAND: [COMMAND].

Used weighted cross-entropy loss to penalise misclassifications of unsafe bash commands. This alleviates the imbalanced dataset.

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 16
  • eval_batch_size: 16
  • seed: 42
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • num_epochs: 5
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss F1 Precision Recall Accuracy
0.1030 1.0 187 0.1318 0.9444 0.9107 0.9808 0.9840
0.1157 2.0 374 0.0823 0.9615 0.9615 0.9615 0.9893
0.1101 3.0 561 0.0543 0.9630 0.9286 1.0 0.9893
0.0812 4.0 748 0.1796 0.9703 1.0 0.9423 0.9920
0.0001 5.0 935 0.0862 0.9524 0.9434 0.9615 0.9866

Framework versions

  • Transformers 5.10.2
  • Pytorch 2.11.0+cu128
  • Datasets 4.0.0
  • Tokenizers 0.22.2
Downloads last month
13
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for P0u4a/ModernBERT-bash-classifier

Finetuned
(1504)
this model