File size: 5,347 Bytes
a7d517c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
---
language:
- en
license: apache-2.0
tags:
- intrusion-detection
- cybersecurity
- tabular
- gradient-boosting
- server-security
- scikit-learn
- dxa-guard
- threat-detection
pipeline_tag: tabular-classification
metrics:
- precision
- recall
- f1
---

# DXA-Guard Intrusion Detection & Prevention Model

Official pre-trained Gradient-Boosted Decision Tree (GBDT) model for **[DXA-Guard](https://github.com/deployxa/dxa-guard)**, an open-source intrusion detection and prevention agent built for Linux VPS servers running web applications.

Published by the **[Deployxa](https://huggingface.co/deployxa)** organization.

This model evaluates **5-minute sliding windows of host and network activity** per source IP to score the likelihood of malicious activity (SSH brute-force, web vulnerability scanners, exploit payloads, and post-compromise behaviors).

---

## Model Architecture & Design

* **Model Type**: Gradient-Boosted Decision Tree Ensemble (`scikit-learn.ensemble.GradientBoostingClassifier`)
* **Number of Trees**: 30 trees (max depth 3)
* **Learning Rate**: 0.1
* **Dual Export Artifacts**:
  1. `model.joblib` / `model.pkl`: Standard scikit-learn serialized models for Python workflows.
  2. `model.json`: Complete portable tree representation containing all decision thresholds and leaf weights. This enables the compiled Go daemon (`dxa-guard`) to evaluate the model natively with **zero Python runtime and sub-millisecond latency**.

### Safety Invariant
The model **never directly triggers bans or kills processes**. In the DXA-Guard architecture, deterministic rules are authoritative for all enforcement actions (nftables bans). This model acts strictly as an advisory risk scorer that elevates or lowers alert urgency without executing actions independently.

---

## Sliding Window Feature Specification (8 Features)

The model takes an 8-dimensional feature vector extracted over a 5-minute sliding window per source IP:

| Index | Feature Name | Type | Description |
|---|---|---|---|
| 0 | `window_event_count` | float | Total number of security events observed for the source IP in the window. |
| 1 | `failure_ratio` | float | Ratio of failed events (failed auth, 401/403/404 HTTP responses) to total events. |
| 2 | `distinct_paths` | float | Count of distinct HTTP request paths targeted by the IP. |
| 3 | `distinct_users` | float | Count of unique usernames attempted during SSH authentication. |
| 4 | `path_entropy` | float | Shannon character entropy of requested URL paths (flags fuzzer / scanner payloads). |
| 5 | `inter_event_time_mean_ms` | float | Mean delta between consecutive events in ms (flags rapid bursts / automated tooling). |
| 6 | `hour_of_day_norm` | float | Timestamp hour divided by 24.0 (0.0 to 1.0) for temporal baseline profiling. |
| 7 | `sensitive_path_ratio` | float | Proportion of requests probing sensitive files (`.env`, `wp-config`, `.git`, webshells). |

---

## Evaluation Benchmark

Evaluated on the DXA-Guard benchmark test set ([`deployxa/dxa-guard-benchmark`](https://huggingface.co/datasets/deployxa/dxa-guard-benchmark)):

| Model / Baseline | Precision | Recall | F1 Score |
|---|:---:|:---:|:---:|
| **DXA-Guard GBDT (This Model)** | **1.000** | **0.833** | **0.909** |
| Isolation Forest (Anomaly Baseline) | 0.941 | 1.000 | 0.970 |

---

## How to Use

### In Python (from Hugging Face)

```python
from huggingface_hub import hf_hub_download
import joblib
import numpy as np

# Download and load model directly from Hugging Face
model_path = hf_hub_download(repo_id="deployxa/dxa-guard", filename="model.joblib")
model = joblib.load(model_path)

# Example feature vector:
# [window_event_count, failure_ratio, distinct_paths, distinct_users,
#  path_entropy, inter_event_time_mean_ms, hour_of_day_norm, sensitive_path_ratio]
sample_features = np.array([[25.0, 0.96, 18.0, 1.0, 4.35, 120.5, 0.58, 0.72]])

# Predict attack probability
prob = model.predict_proba(sample_features)[0, 1]
print(f"Attack Probability: {prob:.4f}")
```

### In Go (Native Inference with zero dependencies)

The model can be evaluated natively in Go by parsing `model.json`:

```go
package main

import (
    "fmt"
    "dxa-guard/internal/scorer"
)

func main() {
    s, err := scorer.LoadModel("model.json")
    if err != nil {
        panic(err)
    }

    features := []float64{25.0, 0.96, 18.0, 1.0, 4.35, 120.5, 0.58, 0.72}
    prob, _ := s.Score(features)
    fmt.Printf("Attack probability: %.4f\n", prob)
}
```

### Native Mathematical Inference Formula

$$logit = \text{init\_bias} + \sum_{t=1}^{30} \text{learning\_rate} \times \text{LeafValue}_t$$
$$P(\text{attack}) = \frac{1}{1 + e^{-logit}}$$

---

## Repository Files

* `model.joblib`: Serialized scikit-learn `GradientBoostingClassifier`.
* `model.pkl`: Standard Python pickle serialized model.
* `model.json`: Native JSON tree export with exact weights and thresholds for Go/C++/Rust runtime.
* `isolation_forest.joblib`: Anomaly detection baseline model.
* `features.py`: Python module containing the reference feature extraction pipeline.
* `features.json`: Machine-readable metadata schema of the 8 feature definitions.
* `test_cases.json`: Reference feature vectors and exact Python probability outputs for cross-language validation.

---

## License & Attribution

Released under the **Apache 2.0 License** by Deployxa.