Text Classification
Transformers
Safetensors
modernbert
Generated from Trainer
text-embeddings-inference
Instructions to use P0u4a/ModernBERT-bash-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use P0u4a/ModernBERT-bash-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="P0u4a/ModernBERT-bash-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("P0u4a/ModernBERT-bash-classifier") model = AutoModelForSequenceClassification.from_pretrained("P0u4a/ModernBERT-bash-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from P0u4a/ModernBERT-bash-classifier: direct link, hf CLI and curl.
- Browser
- Download file 3.24 kB
-
https://huggingface.co/P0u4a/ModernBERT-bash-classifier/resolve/main/README.md
- Command line
-
hf download hf://P0u4a/ModernBERT-bash-classifier/README.md
-
curl -L -o README.md https://huggingface.co/P0u4a/ModernBERT-bash-classifier/resolve/main/README.md
3.24 kB
| library_name: transformers | |
| license: apache-2.0 | |
| base_model: answerdotai/ModernBERT-base | |
| tags: | |
| - generated_from_trainer | |
| metrics: | |
| - f1 | |
| - precision | |
| - recall | |
| - accuracy | |
| model-index: | |
| - name: ModernBERT-bash-classifier | |
| results: [] | |
| # ModernBERT-bash-classifier | |
| This model is a fine-tuned version of [answerdotai/ModernBERT-base](https://huggingface.co/answerdotai/ModernBERT-base) on an unknown dataset. | |
| It achieves the following results on the evaluation set: | |
| - Loss: 0.1192 | |
| - F1: 0.9615 | |
| - Precision: 0.9615 | |
| - Recall: 0.9615 | |
| - Accuracy: 0.9893 | |
| ## Model description | |
| Bash command classification model, for classifying a bash command as either safe or unsafe based on what it does. | |
| We classify a bash command as unsafe based on the following criteria: | |
| - Downloads data from the internet[^1] | |
| - Sends data from the machine over the network[^1] | |
| - Permanently modifies files outside the current project directory[^2] | |
| - Deletes files outside the current project directory[^2] | |
| - Modifies cloud resources (via CLIs like `az` or `tf`) | |
| - Creates persistent, resource consuming processes on the machine like cron jobs | |
| [^1]: Git commands are an exception to this rule. | |
| [^2]: We treat files within the current project directory as safe, since an agent working in a project is expected to mutate its contents. | |
| ## Intended uses & limitations | |
| Created to be used as a auto-mode classifier for coding agent harnesses, so users only manually approve bash commands flagged as unsafe. | |
| The model may not generalize well to bash commands incorporating unknown CLIs or commands that use raw code blocks like Python within heredocs. | |
| ## Training and evaluation data | |
| Fine-tuned on bash tool calls collected from Claude Code and Codex traces. Any PII in the commands such as directory names | |
| were anonymised. The total dataset consists of 3,738 bash calls. | |
| ## Training procedure | |
| Fine-tuned on bash commands formatted as `CWD: [CWD]\nCOMMAND: [COMMAND]`. | |
| Used weighted cross-entropy loss to penalise misclassifications of unsafe bash commands. This alleviates the imbalanced dataset. | |
| ### Training hyperparameters | |
| The following hyperparameters were used during training: | |
| - learning_rate: 2e-05 | |
| - train_batch_size: 16 | |
| - eval_batch_size: 16 | |
| - seed: 42 | |
| - optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments | |
| - lr_scheduler_type: linear | |
| - num_epochs: 5 | |
| - mixed_precision_training: Native AMP | |
| ### Training results | |
| | Training Loss | Epoch | Step | Validation Loss | F1 | Precision | Recall | Accuracy | | |
| |:-------------:|:-----:|:----:|:---------------:|:------:|:---------:|:------:|:--------:| | |
| | 0.1030 | 1.0 | 187 | 0.1318 | 0.9444 | 0.9107 | 0.9808 | 0.9840 | | |
| | 0.1157 | 2.0 | 374 | 0.0823 | 0.9615 | 0.9615 | 0.9615 | 0.9893 | | |
| | 0.1101 | 3.0 | 561 | 0.0543 | 0.9630 | 0.9286 | 1.0 | 0.9893 | | |
| | 0.0812 | 4.0 | 748 | 0.1796 | 0.9703 | 1.0 | 0.9423 | 0.9920 | | |
| | 0.0001 | 5.0 | 935 | 0.0862 | 0.9524 | 0.9434 | 0.9615 | 0.9866 | | |
| ### Framework versions | |
| - Transformers 5.10.2 | |
| - Pytorch 2.11.0+cu128 | |
| - Datasets 4.0.0 | |
| - Tokenizers 0.22.2 | |